SmartOClock: Workload- and Risk-Aware Overclocking in the Cloud
Jovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock, Esha Choukse, Mayukh Das, Chetan Bansal, Jason Lee, Zoey Sun, Haoran Qiu, Reed Zimmermann, Savyasachi Samal
摘要
Operating server components beyond their voltage and power design limits (i.e., overclocking) enables improving performance and lowering cost for cloud workloads. However, overclocking can significantly degrade component lifetime, increase power consumption, and cause power capping events, eventually diminishing the performance benefits.
In this paper, we characterize the impact of overclocking on cloud workloads by studying their profiles from production deployments. Based on the characterization insights, we propose SmartOClock, the first distributed overclocking management platform specifically designed for cloud environments. SmartO-Clock is a workload-aware scheme that relies on power predictions to heterogeneously distribute the power budgets across its servers based on their needs and then enforce budget compliance locally, per-server, in a decentralized manner.
SmartOClock reduces the tail latency by 9%, application cost by 30% and total energy consumption by 10% for latencysensitive microservices on a 36-server deployment. Simulation analysis using production traces show that SmartOClock reduces the number of power capping events by up to 95% while increasing the overclocking success rate by up to 62%. We also describe lessons from building a first-of-its-kind overclockable cluster at a cloud provider for production experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud PlatformsJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse 等ASPLOS 2025 · 被引用 30 次
- Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud PlatformsBenjamin Reidys, Pantea Zardoshti, Íñigo Goiri, Celine Irvene 等ASPLOS 2025 · 被引用 9 次
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo 等HPCA 2026 · 被引用 2 次
- Power Sloshing in Compound Servers for Large-Scale AI Inference WorkloadsAlbert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia 等ISCA 2026 · 被引用 1 次
它引用的顶会 Paper16
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen 等SC 2021 · 被引用 136 次
相关 Paper
- Cost-Efficient Overclocking in Immersion-Cooled DatacentersMajid Jalili, Ioannis Manousakis, Iñigo Goiri, Pulkit A. Misra 等ISCA 2021 · 被引用 54 次
- Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at ScaleShaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis 等OSDI 2020 · 被引用 42 次
- ANT-man: towards agile power management in the microservice eraXiaofeng Hou, Chao Li, Jiacheng Liu, Lu Zhang 等SC 2020 · 被引用 35 次
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang 等NSDI 2025 · 被引用 12 次
