Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving
Junyeol Yu, Jongseok Kim, Euiseong Seo
摘要
The proportion of machine learning (ML) inference in modern cloud workloads is rapidly increasing, and graphic processing units (GPUs) are the most preferred computational accelerators for it. The massively parallel computing capability of GPUs is well-suited to the inference workloads but consumes more power than conventional CPUs. Therefore, GPU servers contribute significantly to the total power consumption of a data center. However, despite their heavy power consumption, GPU power management in cloud-scale has not yet been actively researched. In this paper, we reveal three findings about energy efficiency of ML inference clusters in the cloud. ❶ GPUs of different architectures have comparative advantages in energy efficiency to each other for a set of ML models. ❷ The energy efficiency of a GPU set may significantly vary depending on the number of active GPUs and their clock frequencies even when producing the same level of throughput. ❸ The service level objective(SLO)-blind dynamic voltage and frequency scaling (DVFS) driver of commercial GPUs maintain an immoderately high clock frequency. Based on these implications, we propose a hierarchical GPU resource management approach for cloud-scale inference services. The proposed approach consists of energy-aware cluster allocation, intra-cluster node scaling, intra-node GPU scaling and GPU clock scaling schemes considering the inference service architecture hierarchy. We evaluated our approach with its prototype implementation and cloud-scale simulation. The evaluation with real-world traces showed that the proposed schemes can save up to 28.3% of the cloud-scale energy consumption when serving five ML models with 105 servers having three different kinds of GPUs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Characterizing Power Management Opportunities for LLMs in the CloudPratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri 等ASPLOS 2024 · 被引用 83 次
- Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power SensorZeyu Yang, Karel Adámek, Wesley ArmourSC 2024 · 被引用 38 次
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang 等ASPLOS 2025 · 被引用 9 次
- Power Sloshing in Compound Servers for Large-Scale AI Inference WorkloadsAlbert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia 等ISCA 2026 · 被引用 1 次
相关 Paper
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui 等USENIX ATC 2024 · 被引用 82 次
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu 等EuroSys 2024 · 被引用 18 次
- PowerGrad: Hierarchical Power Management for Power-Limited ML Inference ClustersHyoungwook Nam, Raghavendra Pradyumna Pothukuchi, Alper Buyuktosunoglu, Aporva Amarnath 等ISCA 2026 · 被引用 1 次
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
