The increasing power density of AI infrastructure places critical importance on the thermal behavior of data-center GPUs. Measurements of these GPUs’ temperatures, however, rely on aggregate metrics such as hardware power ratings, data-center efficiency metrics, and node temperatures, none of which provide information about the temperature distribution across the GPUs within that node. As a result, it is important to determine whether thermal divergence among GPU nodes varies with the hardware model alone or with the workloads assigned to those nodes. An analysis of 32 public sessions, each with eight GPUs assigned to either image-generation or language-model and text-generation workloads, shows that the mean temperature spread between the hottest and coolest GPU within each session is 13.58˚C, a value that far exceeds the 1˚C resolution of each GPU’s temperature measurements. Furthermore, analysis of the interaction between workload and hardware reveals that H100 GPUs exhibit a greater temperature spread during image-generation workloads than during language-model and text-generation workloads. In comparison, B200 GPUs exhibit a greater temperature spread during language-model and text-generation workloads than during image-generation workloads (interaction effect: 16.77˚C, F(1, 28) = 67.88, p < 0.001). These results are statistically significant even after adjusting for heteroskedasticity, permutation tests, bootstrap resampling, leave-one-session-out model refitting, and total GPU power within each session. Furthermore, because each set of GPU workload data was collected on different dates, these results indicate only a relationship between workload and GPU thermal divergence, not between the workload itself and that divergence. Thus, thermal assumptions based solely on the hardware model or on the average temperature of GPUs within a node provide incomplete descriptions of the thermal behavior of those nodes. As such, decisions regarding data-center cooling, scheduling workloads across nodes, hardware procurement, and managing the thermal behavior of those GPUs should account for thermal divergence based on each GPU’s workload. Furthermore, while this study was unable to provide evidence of any effect of these workloads on the throttling or useful computation of those GPUs, such conclusions would require validation of these assumptions.
International Energy Agency (2025) Energy and AI (World Energy Outlook Special Report). IEA. https://www.iea.org/reports/energy-and-ai
Lawrence Berkeley National Laboratory (2024) 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/publications/2024-lbnl-data-center-energy-usage-report
Electric Power Research Institute (2026) Powering Intelligence 2026: Updated Scenarios of U.S. Data Center Electricity Use and Power Strategies (Report No. 3002034696). EPRI. https://www.epri.com/research/products/000000003002034696
ASHRAE, NEMA and Pacific Northwest National Laboratory (2026) AI Data Center Energy Performance Framework. ASHRAE, NEMA and PNNL. https://www.ashrae.org/technical-resources/ai-data-center-framework
The Green Grid (2012) PUE: A Comprehensive Examination of the Metric (White Paper No. 49; V. Avelar, D. Azevedo, & A. French, Eds.). The Green Grid. https://archive.thegreengrid.org/en/resources/library-and-tools/237-WP
Horner, N. and Azevedo, I. (2016) Power Usage Effectiveness in Data Centers: Overloaded and Underachieving. The Electricity Journal , 29, 61-69. https://doi.org/10.1016/j.tej.2016.04.011
Fan, X., Weber, W. and Barroso, L.A. (2007) Power Provisioning for a Warehouse-Sized Computer. Proceedings of the 34 th Annual International Symposium on Computer Architecture , San Diego, 9-13 June 2007, 13-23. https://doi.org/10.1145/1250662.1250665
Skadron, K., Stan, M.R., Huang, W., Velusamy, S., Sankaranarayanan, K. and Tarjan, D. (2003) Temperature-Aware Microarchitecture. Proceedings of the 30 th Annual International Symposium on Computer Architecture—ISCA ’03, San Diego, 9-11 June 2003, 2-13. https://doi.org/10.1145/859618.859620
Huang, W., Ghosh, S., Velusamy, S., Sankaranarayanan, K., Skadron, K. and Stan, M.R. (2006) Hotspot: A Compact Thermal Modeling Methodology for Early-Stage VLSI Design. IEEE Transactions on Very Large Scale Integration ( VLSI ) Systems , 14, 501-513. https://doi.org/10.1109/tvlsi.2006.876103
Hamann, H.F., Weger, A., Lacey, J.A., Hu, Z., Bose, P., Cohen, E., et al . (2007) Hotspot-Limited Microprocessors: Direct Temperature and Power Distribution Measurements. IEEE Journal of Solid-State Circuits , 42, 56-65. https://doi.org/10.1109/jssc.2006.885064
Patel, P., Choukse, E., Zhang, C., Goiri, Í., Warrier, B., Mahalingam, N., et al . (2024) Characterizing Power Management Opportunities for LLMs in the Cloud. Proceed ings of the 29 th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , Volume 3, La Jolla, 27 April-1 May 2024, 207-222. https://doi.org/10.1145/3620666.3651329
Mayr, M., Wind, S., Schroder, L., Moradi, M., Hager, G., Kostler, H. and Wellein, G. (2026) AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models. arXiv: 2603.16164. https://arxiv.org/abs/2603.16164
Go, S., Park, J., More, S., Wu, H., Wang, I., Jezghani, A., et al . (2025) Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective. Proceedings of the 58 th IEEE / ACM International Symposium on Microarchitecture , Seoul, 18-22 October 2025, 626-642. https://doi.org/10.1145/3725843.3756111
Chung, J.W., Ma, J.J., Wu, R., Liu, J., Kweon, O.J., Xia, Y., Wu, Z. and Chowdhury, M. (2025) The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization. arXiv: 2505.06371. https://arxiv.org/abs/2505.06371
Špeťko, M., Vysocký, O., Jansík, B. and Říha, L. (2021) DGX-A100 Face to Face Dgx-2—Performance, Power and Thermal Behavior Evaluation. Energies , 14, Article 376. https://doi.org/10.3390/en14020376
Elsayed, A.A.E., Al-Obaidi, A.A. and Farag, H.E.Z. (2026) Characterization of High-Resolution AI Data Center Training Workloads on Single and Multiple GPU Nodes. Scientific Data . https://doi.org/10.1038/s41597-026-07496-6
Masanet, E., Shehabi, A., Lei, N., Smith, S. and Koomey, J. (2020) Recalibrating Global Data Center Energy-Use Estimates. Science , 367, 984-986. https://doi.org/10.1126/science.aba3758
Kuzay, M., Demirel, E., Bayraktar, B., Vilestad, J., Kärnebro, A., Yilmaz, C., et al . (2026) A Self-Assessment Framework for Evaluating Efficiency of Data Centers. Energy Informatics , 9, Article No. 39. https://doi.org/10.1186/s42162-026-00652-7
You, J., Chung, J.W. and Chowdhury, M. (2023) Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training. Proceedings of the 20 th USENIX Symposium on Networked Systems Design and Implementation ( NSDI ’23), Boston, 17-19 April 2023, 119-139. https://www.usenix.org/conference/nsdi23/presentation/you
Tschand, A., Rajan, A.T.R., Idgunji, S., Ghosh, A., Holleman, J., Kiraly, C., et al . (2025) MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from μ Watts to MWatts for Sustainable AI. 2025 IEEE International Symposium on High Performance Computer Architecture ( HPCA ), Las Vegas, 1-5 March 2025, 1201-1216. https://doi.org/10.1109/hpca61900.2025.00092
Fadel Argerich, M., Furst, J. and Patino-Martinez, M. (2026) Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures. arXiv: 2604.09048. https://arxiv.org/abs/2604.09048
Hong, S. and Kim, H. (2010) An Integrated GPU Power and Performance Model. Proceedings of the 37 th Annual International Symposium on Computer Architecture , Saint-Malo, 19-23 June 2010, 280-289. https://doi.org/10.1145/1815961.1815998
Leng, J., Hetherington, T., ElTantawy, A., Gilani, S., Kim, N.S., Aamodt, T.M., et al . (2013) GPUWattch: Enabling Energy Optimizations in GPGPUs. Proceedings of the 40 th Annual International Symposium on Computer Architecture , Tel-Aviv, 23-27 June 2013, 487-498. https://doi.org/10.1145/2485922.2485964
Kandiah, V., Peverelle, S., Khairy, M., Pan, J., Manjunath, A., Rogers, T.G., et al . (2021) AccelWattch: A Power Modeling Framework for Modern GPUs. MICRO -54: 54 th Annual IEEE / ACM International Symposium on Microarchitecture , 18-22 October 2021, 738-753. https://doi.org/10.1145/3466752.3480063
NVIDIA (n.d.) Data Center GPU Manager Documentation. https://docs.nvidia.com/datacenter/dcgm/
Yang, Z., Adamek, K. and Armour, W. (2024) Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU’s Built-In Power Sensor. SC 24: International Conference for High Performance Computing , Networking , Storage and Analysis , Atlanta, 17-22 November 2024, 1-17. https://doi.org/10.1109/sc41406.2024.00028
Garimella, S.V., Fleischer, A.S., Murthy, J.Y., Keshavarzi, A., Prasher, R., Patel, C., et al . (2008) Thermal Challenges in Next-Generation Electronic Systems. IEEE Transactions on Components and Packaging Technologies , 31, 801-815. https://doi.org/10.1109/tcapt.2008.2001197
Szekely, V. (1998) Identification of RC Networks by Deconvolution: Chances and Limits. IEEE Transactions on Circuits and Systems I : Fundamental Theory and Applications , 45, 244-258. https://doi.org/10.1109/81.662698
Lasance, C.J.M. (2008) Ten Years of Boundary-Condition-Independent Compact Thermal Modeling of Electronic Parts: A Review. Heat Transfer Engineering , 29, 149-168. https://doi.org/10.1080/01457630701673188
JEDEC Solid State Technology Association (2010) Transient Dual Interface Test Method for the Measurement of the Thermal Resistance Junction-to-Case of Semi-Conductor Devices with Heat Flow through a Single Path (JESD51-14). JEDEC.
Farkas, G., Poppe, A. and Rencz, M. (2022) Theoretical Background of Thermal Transient Measurements. In: Rencz, M., Farkas, G. and Poppe, A., Eds., Theory and Practice of Thermal Transient Testing of Electronic Components , Springer, 7-96. https://doi.org/10.1007/978-3-030-86174-2_2
Sridhar, A., Vincenzi, A., Ruggiero, M., Brunschwiler, T. and Atienza, D. (2010) 3D-ICE: Fast Compact Transient Thermal Modeling for 3D ICs with Inter-Tier Liquid Cooling. 2010 IEEE / ACM International Conference on Computer - Aided Design ( ICCAD ), San Jose, 7-11 November 2010, 463-470. https://doi.org/10.1109/iccad.2010.5653749
Price, D.C., Clark, M.A., Barsdell, B.R., Babich, R. and Greenhill, L.J. (2015) Optimizing Performance-per-Watt on GPUs in High Performance Computing. Computer Science — Research and Development , 31, 185-193. https://doi.org/10.1007/s00450-015-0300-5
Jia, Z., Maggioni, M., Smith, J. and Scarpazza, D.P. (2019) Dissecting the NVIDIA Turing T4 GPU via Microbenchmarking. arXiv: 1903.07486. https://arxiv.org/abs/1903.07486
Dean, J. and Barroso, L.A. (2013) The Tail at Scale. Communications of the ACM , 56, 74-80. https://doi.org/10.1145/2408776.2408794
Chen, J., Pan, X., Monga, R., Bengio, S. and Jozefowicz, R. (2016) Revisiting Distributed Synchronous SGD. arXiv: 1604.00981. https://arxiv.org/abs/1604.00981
Lin, J., Jiang, Z., Song, Z., Zhao, S., Yu, M., Wang, Z., Wang, C., Shi, Z., Shi, X., Jia, W., Liu, Z., Wang, S., Lin, H., Liu, X., Panda, A. and Li, J. (2025) Understanding Stragglers in Large Model Training Using What-If Analysis. Proceedings of the 19 th USENIX Symposium on Operating Systems Design and Implementation ( OSDI ’25), Boston, 7-9 July 2025, 483-498. https://www.usenix.org/conference/osdi25/presentation/lin-jinkun
Moore, J., Chase, J., Ranganathan, P. and Sharma, R. (2005) Making Scheduling Cool: Temperature-Aware Workload Placement in Data Centers. Proceedings of the 2005 USENIX Annual Technical Conference , Anaheim, 10-15 April 2005, 61-75. https://www.usenix.org/conference/2005-usenix-annual-technical-conference/making-scheduling-cool-temperature-aware-workload
Stojkovic, J., Zhang, C., Goiri, Í., Choukse, E., Qiu, H., Fonseca, R., et al . (2025) TAPAS: Thermal-and Power-Aware Scheduling for LLM Inference in Cloud Platforms. Proceedings of the 30 th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , Volume 2, Rotterdam, 30 March-3 April 2025 1266-1281. https://doi.org/10.1145/3676641.3716025
Lu, R. and Wang, D. (2025) A Thermal-Aware Workload Scheduler for High-Performance LLM Inference in Cooling-Regulated Datacenters. ACM SIGEnergy Energy Informatics Review , 5, 98-104. https://doi.org/10.1145/3757892.3757906
Open Compute Project (n.d.) ACS Liquid Cooling Cold Plate Requirements. Open Compute Project. https://www.opencompute.org/documents/ocp-acs-liquid-cooling-cold-plate-requirements-pdf
Open Compute Project (n.d.) OAI System Liquid Cooling Guidelines. Open Compute Project. https://www.opencompute.org/documents/oai-system-liquid-cooling-guidelines-in-ocp-template-mar-3-2023-update-pdf
ARPA-E (n.d.) COOLERCHIPS Program. U.S. Department of Energy. https://arpa-e.energy.gov/technologies/programs/coolerchips
Bhatasana, M. and Marconnet, A.M. (2025) Phase Change Materials as Thermal Buffers for Power Electronics Modules with Transient Heat Loads. Energy Conversion and Management , 343, Article ID: 119931. https://doi.org/10.1016/j.enconman.2025.119931
Open Compute Project (n.d.) Cooling Environments. Open Compute Project. https://www.opencompute.org/wiki/Cooling_Environments