PowerGrad: Hierarchical Power Management for Power-Limited ML Inference Clusters
Hyoungwook Nam, Raghavendra Pradyumna Pothukuchi, Alper Buyuktosunoglu, Aporva Amarnath, Pradip Bose, Josep Torrellas
摘要
As machine learning (ML) workloads demand more power and datacenters integrate renewable energy, workloads have to deal with situations where power demands exceed supply. In such situations, intelligently allocating the power among the nodes is key to maximizing efficiency. However, this is hard to do for ML inference workloads, where system administrators cannot profile the workload ahead of time. To address this challenge, this paper proposes PowerGrad, a hierarchical power-management framework for power-limited ML inference clusters. The idea is to dynamically identify the performance gradient of each running workload, which characterizes the performance sensitivity of the workload to power changes. At runtime, a Gradient Estimator collects hardware measurements and uses them to estimate performance gradients. Then, to maximize efficiency, Local Controllers and Hierarchical Controllers re-distribute the power from low-gradient workloads to high-gradient ones within a node and across nodes, respectively. PowerGrad is especially effective for severely power-limited environments, where every node demands more power than its maximum allocation. While PowerGrad can be applied to a variety of compute architectures, it needs dynamic hardware performance counter information that is unavailable in GPUs and accelerators. Consequently, we demonstrate PowerGrad on two CPU clusters running popular ML inference workloads in power-limited setups. The results show that PowerGrad is both effective and easily retargetable across different architectures. In traditional dualCPU nodes, PowerGrad reduces the average and tail latencies by a mean of and , respectively, relative to the strongest of a set of software-transparent baselines. In single-CPU nodes with ML acceleration support, PowerGrad reduces the average and tail latencies by a mean of 9.0% and 9.9%, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 被引用 14 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair 等SC 2024 · 被引用 10 次
- Themis: Fair and Efficient GPU Cluster SchedulingKshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman 等NSDI 2020 · 被引用 22 次
