Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads
Albert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia, Sultan Mahmud Sajal, Gefei Zuo, Krishna T. Malladi, Devon Akers, Kalyan Subramanian, Shobhit O. Kanaujia, Alexandros Daglis
Abstract
AI workloads are rapidly emerging as an important component of datacenter operations, with inference services in particular consuming an ever-increasing share of computational cycles. To keep pace with the growing demand for AI-driven applications, datacenters have begun deploying advanced compound servers that tightly integrate accelerators such as GPUs with traditional CPUs. While these platforms enable high performance, they also drive up the power requirements of datacenters, creating new challenges for scaling the infrastructure. To address these challenges, we conduct a comprehensive characterization of power usage patterns in datacenters hosting AI services. We find that power consumption fluctuates widely across time, models, services, and server components. These observations reveal the inefficiency of applying fixed, uniform power limits across servers, which can result in either under-provisioning and degraded performance, or overprovisioning and wasted planned power. Motivated by these insights, we investigate dynamic power control to optimize both power consumption and performance. Through experiments on production workloads, we demonstrate that controlled power sloshing can deliver benefits, yielding up to 30% power savings. We further develop a scalable, automated algorithm for server-level power management, which could be deployed in datacenters and reduce power by up to 11% without degrading the quality of service. Finally, based on our production experience, we offer practical guidelines for hardware and software co-design in next-generation AI platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de651467-9cfb-400c-803c-2ad057fa7e4aBuilds on23
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 220 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP InferenceThierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia et al.MICRO 2021 · 117 citations
Related papers
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 14 citations
- Characterizing Power Management Opportunities for LLMs in the CloudPratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri et al.ASPLOS 2024 · 83 citations
- PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power ManagementVasilis Kypriotis, Eric Dubberstein, Patrick H. Coppock, Eliot H. Solomon et al.ISCA 2026 · 1 citation
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui et al.USENIX ATC 2024 · 82 citations
- PowerGrad: Hierarchical Power Management for Power-Limited ML Inference ClustersHyoungwook Nam, Raghavendra Pradyumna Pothukuchi, Alper Buyuktosunoglu, Aporva Amarnath et al.ISCA 2026 · 1 citation
