Minimizing Power Waste in Heterogenous Computing via Adaptive Uncore Scaling
Zhong Zheng, Seyfal Sultanov, Michael E. Papka, Zhiling Lan
摘要
High-performance computing (HPC) systems are essential for scientific discovery and engineering innovation. However, their growing power demands pose significant challenges, particularly as systems scale to the exascale level. Prior uncore frequency tuning studies have primarily focused on conventional HPC workloads running on CPU-only systems. As HPC advances toward heterogeneous computing, integrating diverse GPU workloads on heterogeneous CPU-GPU systems, it becomes imperative to revisit and enhance uncore scaling. Our investigation reveals that uncore frequency scales down only when CPU power approaches its thermal design power (TDP), which is rare in GPU-dominant applications. As a result, modern computing systems experience unnecessary power waste. In this study, we present MAGUS, a user-transparent uncore frequency scaling runtime for heterogeneous computing. MAGUS dynamically adjusts uncore frequencies according to distinct application execution phases, effectively minimizing power waste caused by consistently using maximum uncore frequencies. Our design incorporates several key techniques, including real-time monitoring and prediction of memory accesses, intelligent handling of frequent phase transitions, and leveraging vendor-provided power management features. We evaluate MAGUS with various GPU benchmarks and applications on multiple heterogeneous systems with different CPU and GPU architectures. Experimental results demonstrate that MAGUS achieves up to 27% energy savings compared to the default settings, while maintaining a performance loss of less than 5% and an overhead of under 1%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau 等USENIX ATC 2020 · 被引用 398 次
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 被引用 220 次
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan 等MICRO 2021 · 被引用 134 次
- EnvPipe: Performance-preserving DNN Training Framework for Saving EnergySangjin Choi, Inhoe Koo, Jeongseob Ahn, Myeongjae Jeon 等USENIX ATC 2023 · 被引用 41 次
- DPS: Adaptive Power Management for Overprovisioned SystemsJianru Ding, Henry HoffmannSC 2023 · 被引用 8 次
相关 Paper
- EVeREST: An Effective and Versatile Runtime Energy Saving Tool for GPUsAnna Yue, Pen-Chung Yew, Sanyam MehtaPPoPP 2025 · 被引用 3 次
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui 等USENIX ATC 2024 · 被引用 82 次
- SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy SavingKaijie Fan, Marco D'Antonio, Lorenzo Carpentieri, Biagio Cosenza 等SC 2023 · 被引用 13 次
- Cuttlefish: library for achieving energy efficiency in multicore parallel programsSunil Kumar, Akshat Gupta, Vivek Kumar, Sridutt BhalachandraSC 2021 · 被引用 9 次
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu 等EuroSys 2024 · 被引用 18 次
