Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance Assurance
Yijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu, Bingqiang Wang
摘要
Power consumption is one of the top limiting factors in high-performance computing systems and data centers, and dynamic voltage and frequency scaling (DVFS) is an important mechanism to control power. Existing works using DVFS to improve GPU energy efficiency suffer from the limitation that their policies either impact performance too much or require offline application profiling or code modification, which severely limits their applicability on large clusters. To address this issue, we propose a novel GPU DVFS policy, GEEPAFS, which improves the energy efficiency of GPUs while providing performance assurance. GEEPAFS is application-transparent as it does not require any offline profiling or code modification on user applications. To achieve this, GEEPAFS models application performance online based on our quantitative analysis of a correlation between performance and GPU memory bandwidth utilization. Based on their relationship, GEEPAFS builds a fold-line frequency-performance model for applications being executed, and it applies the model to guide the setting of GPU frequency to maximize energy efficiency while ensuring the performance loss is bounded. Through experiments on NVIDIA V100 and A100 GPUs, we show that GEEPAFS is able to improve the energy efficiency by 26.7% and 20.2% on average. While achieving this improvement, the average performance loss is only 5.8%, and the worst-case performance loss is 12.5% among all 33 tested applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang 等ASPLOS 2025 · 被引用 9 次
- LithOS: An Operating System for Efficient Machine Learning on GPUsPatrick H. Coppock, Brian Zhang, Eliot H. Solomon, Vasilis Kypriotis 等SOSP 2025 · 被引用 4 次
- Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared SettingsPierre Jacquet, Maxime Agusti, Eddy Caron, Camille Coti 等EuroSys 2026 · 被引用 2 次
- Power Sloshing in Compound Servers for Large-Scale AI Inference WorkloadsAlbert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia 等ISCA 2026 · 被引用 1 次
相关 Paper
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 被引用 14 次
- Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUsSrikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann 等ASPLOS 2023 · 被引用 15 次
- PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power ManagementVasilis Kypriotis, Eric Dubberstein, Patrick H. Coppock, Eliot H. Solomon 等ISCA 2026 · 被引用 1 次
- GVProf: a value profiler for GPU-based clustersKeren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng 等SC 2020 · 被引用 24 次
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 被引用 13 次
