Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at Scale
Shaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis, Sreekumar Kodakara, David Lo, Parthasarathy Ranganathan
摘要
As the demand for data center capacity continues to grow, hyperscale providers have used power oversubscription to increase efficiency and reduce costs. Power oversubscription requires power capping systems to smooth out the spikes that risk overloading power equipment by throttling workloads. Modern compute clusters run latency-sensitive serving and throughput-oriented batch workloads on the same servers, provisioning resources to ensure low latency for the former while using the latter to achieve high server utilization. When power capping occurs, it is desirable to maintain low latency for serving tasks and throttle the throughput of batch tasks. To achieve this, we seek a system that can gracefully throttle batch workloads and has task-level quality-of-service (QoS) differentiation.
In this paper we present Thunderbolt, a hardware-agnostic power capping system that ensures safe power oversubscription while minimizing impact on both long-running throughput-oriented tasks and latency-sensitive tasks. It uses a two-threshold, randomized unthrottling/multiplicative decrease control policy to ensure power safety with minimized performance degradation. It leverages the Linux kernel's CPU bandwidth control feature to achieve task-level QoS-aware throttling. It is robust even in the face of power telemetry unavailability. Evaluation results at the node and cluster levels demonstrate the system's responsiveness, effectiveness for reducing power, capability of QoS differentiation, and minimal impact on latency and task health. We have deployed this system at scale, in multiple production clusters. As a result, we enabled power oversubscription gains of 9%-25%, where none was previously possible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Characterizing Power Management Opportunities for LLMs in the CloudPratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri 等ASPLOS 2024 · 被引用 83 次
- Ecovisor: A Virtual Energy System for Carbon-Efficient ApplicationsAbel Souza, Noman Bashir, Jorge Murillo, Walid A. Hanafy 等ASPLOS 2023 · 被引用 56 次
- Flex: High-Availability Datacenters With Zero Reserved PowerChaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang 等ISCA 2021 · 被引用 43 次
它引用的顶会 Paper2
相关 Paper
- PASS: A Power Adaptive Storage ServerDedong Xie, Theano Stavrinos, Jonggyu Park, Simon Peter 等EuroSys 2026
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock 等ISCA 2024 · 被引用 12 次
- ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the CloudShuang Chen, Angela Jin, Christina Delimitrou, José F. MartínezHPCA 2022 · 被引用 34 次
- Prediction-Informed Power Management for General-Purpose Compute ServersJonggyu Park, Simon Peter, Thomas E. AndersonEuroSys 2026
