Article citationsMore>>
Lin, J., Jiang, Z., Song, Z., Zhao, S., Yu, M., Wang, Z., Wang, C., Shi, Z., Shi, X., Jia, W., Liu, Z., Wang, S., Lin, H., Liu, X., Panda, A. and Li, J. (2025) Understanding Stragglers in Large Model Training Using What-If Analysis. Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI’25), Boston, 7-9 July 2025, 483-498.
https://www.usenix.org/conference/osdi25/presentation/lin-jinkun
has been cited by the following article:
-
TITLE:
Workload-Dependent Thermal Locality in AI GPU Nodes: Evidence from H100 and B200 Intra-Node Thermal Divergence
AUTHORS:
Daniel Weon Seok Ko, Johnathan Mun
KEYWORDS:
AI Infrastructure, H100, B200, Thermal Locality, Node Thermal Spread, Workload Measurement, Data-Center Cooling, Thermal Management, GPU Heat Management
JOURNAL NAME:
Intelligent Control and Automation,
Vol.17 No.3,
July
24,
2026
ABSTRACT: The increasing power density of AI infrastructure places critical importance on the thermal behavior of data-center GPUs. Measurements of these GPUs’ temperatures, however, rely on aggregate metrics such as hardware power ratings, data-center efficiency metrics, and node temperatures, none of which provide information about the temperature distribution across the GPUs within that node. As a result, it is important to determine whether thermal divergence among GPU nodes varies with the hardware model alone or with the workloads assigned to those nodes. An analysis of 32 public sessions, each with eight GPUs assigned to either image-generation or language-model and text-generation workloads, shows that the mean temperature spread between the hottest and coolest GPU within each session is 13.58˚C, a value that far exceeds the 1˚C resolution of each GPU’s temperature measurements. Furthermore, analysis of the interaction between workload and hardware reveals that H100 GPUs exhibit a greater temperature spread during image-generation workloads than during language-model and text-generation workloads. In comparison, B200 GPUs exhibit a greater temperature spread during language-model and text-generation workloads than during image-generation workloads (interaction effect: 16.77˚C, F(1, 28) = 67.88, p