ReThermal: Co-Design of Thermal-Aware Static and Dynamic Scheduling for LLM Training on Liquid-Cooled Wafer-Scale Chips
Chengran Li, Huizheng Wang, Jiaxin Liu, Jingyao Liu, Zhiheng Yue, Xia Li, Shenfei Jiang, Jinyi Deng, Yang Hu, Shouyi Yin
Abstract
With the increasing demand for high computational power in Large Language Models, wafer-scale chips have emerged as a solution, providing the necessary integration and computing capability to meet these needs. However, their ultralarge area and extreme heat dissipation introduce critical thermal management challenges under liquid-cooling environments. In addressing this issue, we identify two key opportunities and three major challenges: the behavior-thermal black box, the waferscale simulation bottleneck, and the runtime heat-schedule drift. To tackle these challenges, we propose ReThermal, a holistic scheduling framework that integrates three innovations. First, we introduce behavior-driven thermal modeling to capture workload-induced compute, communication, and heat coupling patterns at the system level. Second, we develop a DNNaccelerated wafer-scale thermal simulator that enables fast and accurate temperature prediction, significantly reducing simulation time. Third, we implement an adaptive thermal-aware scheduling strategy that coordinates compile-time and runtime decisions to dynamically optimize task placement. Evaluations show that ReThermal reduces peak temperature by up to 8.0° C and improves throughput by up to 39.23 %, providing a scalable and effective thermal control solution for future liquid-cooled wafer-scale systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e795b2f9-190e-4554-8380-18f4c39ab1a2Related papers
- WSC-LLM: Efficient LLM Service and Architecture Co-exploration for Wafer-scale ChipsZheng Xu, Dehao Kong, Jiaxin Liu, Jinxi Li et al.ISCA 2025 · 22 citations
- FACE: Fully Overlapped PD Scheduling and Multi-Level Architecture Co-Exploration on WaferZheng Xu, Dehao Kong, Jiaxin Liu, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou et al.HPCA 2026 · 2 citations
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 11 citations
