Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
Ruofan Wu, Jae-Won Chung, Mosharaf Chowdhury
摘要
The computing demand of AI is growing at an unprecedented rate, but energy supply is not keeping pace. As a result, energy has become an expensive and contended resource that requires explicit management and optimization. Although recent works have made significant progress in large model training optimization, they focus on optimizing either dynamic or static energy consumption.
We find that fine-grained kernel scheduling and frequency scaling jointly and interdependently impact both dynamic and static energy consumption. Based on this finding, we design Kareus, a training system that pushes the time-energy tradeoff frontier by optimizing both aspects. Kareus decomposes the intractable joint optimization problem into local, partitionbased subproblems. It then uses a multi-pass multi-objective optimization algorithm to find execution schedules that push the time-energy tradeoff frontier. Compared to the state of the art, Kareus reduces training energy by up to 28.3% at the same training time, or reduces training time by up to 27.5% at the same energy consumption. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 被引用 220 次
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan 等MICRO 2021 · 被引用 134 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- NanoFlow: Towards Optimal Large Language Model Serving ThroughputKan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao 等OSDI 2025 · 被引用 92 次
相关 Paper
- Reducing Energy Bloat in Large Model TrainingJae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng 等SOSP 2024 · 被引用 12 次
- EA-HAS-Bench: Energy-aware Hyperparameter and Architecture Search BenchmarkShuguang Dou, Xinyang Jiang, Cairong Zhao, Dongsheng LiICLR 2023
- SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy SavingKaijie Fan, Marco D'Antonio, Lorenzo Carpentieri, Biagio Cosenza 等SC 2023 · 被引用 13 次
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang 等ASPLOS 2025 · 被引用 9 次
- FlipFlop: A Static Analysis-based Energy Optimization Framework for GPU KernelsSaurabhsingh Rajput, Alexander Brandt, Vadim Elisseev, Tushar SharmaICSE 2026
