Flex: High-Availability Datacenters With Zero Reserved Power
Chaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang, Pulkit A. Misra, Rod Assis, Kyle Woolcock, Nithish Mahalingam, Brijesh Warrier, David Gauthier, Lalu Kunnath, Steve Solomon
摘要
Cloud providers, like Amazon and Microsoft, must guarantee high availability for a large fraction of their workloads. For this reason, they build datacenters with redundant infrastructures for power delivery and cooling. Typically, the redundant resources are reserved for use only during infrastructure failure or maintenance events, so that workload performance and availability do not suffer. Unfortunately, the reserved resources also produce lower power utilization and, consequently, require more datacenters to be built. To address these problems, in this paper we propose "zero-reserved-power" datacenters and the Flex system to ensure that workloads still receive their desired performance and availability. Flex leverages the existence of software-redundant workloads that can tolerate lower infrastructure availability, while imposing minimal (if any) performance degradation for those that require high infrastructure availability. Flex mainly comprises (1) a new offline workload placement policy that reduces stranded power while ensuring safety during failure or maintenance events, and (2) a distributed system that monitors for failures and quickly reduces the power draw while respecting the workloads’ requirements, when it detects a failure. Our evaluation shows that Flex produces less than 5% stranded power and increases the number of deployed servers by up to 33%, which translates to hundreds of millions of dollars in construction cost savings per datacenter site. We end the paper with lessons from our experience bringing Flex to production in Microsoft’s datacenters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Carbon Explorer: A Holistic Framework for Designing Carbon Aware DatacentersBilge Acun, Benjamin C. Lee, Fiodar Kazhamiaka, Kiwan Maeng 等ASPLOS 2023 · 被引用 171 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Characterizing Power Management Opportunities for LLMs in the CloudPratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri 等ASPLOS 2024 · 被引用 83 次
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene 等ISCA 2024 · 被引用 49 次
- TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud PlatformsJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse 等ASPLOS 2025 · 被引用 30 次
它引用的顶会 Paper3
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
- Data Center Power Oversubscription with a Medium Voltage Power Plane and Priority-Aware CappingVarun Sakalkar, Vasileios Kontorinis, David Landhuis, Shaohong Li 等ASPLOS 2020 · 被引用 56 次
- Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at ScaleShaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis 等OSDI 2020 · 被引用 42 次
相关 Paper
- Hyrax: Fail-in-Place Server Operation in Cloud PlatformsJialun Lyu, Marisa You, Celine Irvene, Mark Jung 等OSDI 2023 · 被引用 18 次
- FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical BackbonesCongcong Miao, Zhizhen Zhong, Ying Zhang, Kunling He 等SIGCOMM 2023 · 被引用 10 次
- You Can Always Get What You Want: CPU Virtualization Made Fast and FreeYun Wang, Xingguo Jia, Ben Luo, Kenan Liu 等SOSP 2026
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- SmartHarvest: harvesting idle CPUs safely and efficiently in the cloudYawen Wang, Kapil Arya, Marios Kogias, Manohar Vanga 等EuroSys 2021 · 被引用 53 次
