Benchmark Overview
GPUs throttle clock speeds when temperatures exceed safe thresholds, silently reducing AI inference throughput. A well-cooled RTX 6000 Pro running at 1,410 MHz delivers 48 tok/s. The same GPU throttled to 1,200 MHz at 85C delivers 38 tok/s, a 21% performance loss with no visible error. We measured thermal throttling behavior across GPU models during sustained AI workloads on dedicated GPU hosting.
Test Configuration
GPUs: RTX 5090, RTX 6000 Pro, RTX 6000 Pro 96 GB, RTX 6000 Pro. Workload: continuous Llama 3 70B INT4 inference via vLLM at 10 concurrent users for 60 minutes. Temperature and clock speed logged every second via nvidia-smi. Ambient temperature: 22C (data centre baseline). See token benchmarks for baseline throughput numbers.
Temperature and Performance Over Time
| GPU | Temp @ 5min | Temp @ 30min | Temp @ 60min | Clock Drop | Throughput Loss |
|---|---|---|---|---|---|
| RTX 5090 (air) | 68C | 78C | 82C | 2,520 to 2,100 MHz | -15% |
| RTX 6000 Pro (server) | 62C | 70C | 72C | 2,520 to 2,400 MHz | -5% |
| RTX 6000 Pro (server) | 55C | 65C | 68C | 1,410 to 1,380 MHz | -2% |
| RTX 6000 Pro (server) | 58C | 68C | 72C | 1,980 to 1,920 MHz | -3% |
Throttling Thresholds
| GPU | Throttle Start | Severe Throttle | Shutdown Temp |
|---|---|---|---|
| RTX 5090 | 75C | 83C | 90C |
| RTX 6000 Pro | 78C | 85C | 92C |
| RTX 6000 Pro | 80C | 85C | 92C |
| RTX 6000 Pro | 80C | 85C | 92C |
Analysis
The RTX 5090 shows the most throttling because its consumer-grade air cooling cannot dissipate 380W of sustained heat as effectively as server-grade cooling in data centre GPUs. The RTX 6000 Pro runs coolest due to its lower 300W TDP and enterprise passive cooling designed for sustained workloads. Check GPU comparisons for thermal specifications.
Data centre environments with proper airflow keep RTX 6000 Pro and RTX 6000 Pro GPUs well below throttling thresholds. Home or office deployments of RTX 5090 may see 10-15% throughput degradation after 30 minutes of sustained inference, particularly in multi-GPU configurations where GPUs share airspace. See multi-GPU cluster cooling requirements.
Cooling Strategies
For private AI hosting in data centres, server-grade cooling eliminates throttling for RTX 6000 Pro and RTX 6000 Pro GPUs. For RTX 5090 in desktop chassis, ensure adequate case airflow and consider aftermarket cooling. In multi-GPU setups, space GPUs at least one slot apart. For managed dedicated servers, cooling is handled by the provider. Configure monitoring in your vLLM production deployment.
Recommendations
Monitor GPU temperatures during sustained inference. If temperatures exceed 75C consistently, throughput is silently degraded. Data centre GPUs (RTX 6000 Pro, RTX 6000 Pro, RTX 6000 Pro) with proper cooling maintain near-peak performance indefinitely. Consumer GPUs (RTX 5090) in non-data-centre environments should budget 10-15% throughput reduction. Deploy on GigaGPU dedicated servers with professional cooling. See the benchmarks section, LLM hosting guide, and infrastructure blog for more data.