Hyrax: Fail-in-Place Server Operation in Cloud Platforms
Jialun Lyu, Marisa You, Celine Irvene, Mark Jung, Tyler Narmore, Jacob Shapiro, Luke Marshall, Savyasachi Samal, Ioannis Manousakis, Lisa Hsu, Preetha Subbarayalu, Ashish Raniwala
摘要
Today's cloud platforms handle server hardware failures by shutting down the affected server and only turning it back online once it has been repaired by a technician. At cloud scale, this all-or-nothing operating model is becoming increasingly unsustainable. This model is also at odds with technology trends, such as the need for new cooling technology.
This paper introduces Hyrax, a datacenter stack that enables compute servers with failed components to continue hosting VMs while hiding the underlying degraded capacity and performance. A key enabler of Hyrax is a novel model of changes in memory interleaving when deactivating faulty memory modules. Experiments on cloud production servers show that Hyrax overcomes common hardware failures without impacting peak VM performance. In large-scale simulations with production traces, Hyrax reduces server repair requirements by 50-60% without impacting VM scheduling.
- Formerly at Microsoft Azure 1 Prior work reported 0.09% [58, 59], 0.12% [12], and 1.6% [57,66] for DIMMs and 0.22% [43,44] to 1.2% [3,54] for SSDs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Understanding Silent Data Corruptions in a Large Production CPU PopulationShaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang 等SOSP 2023 · 被引用 56 次
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene 等ISCA 2024 · 被引用 49 次
- FairyWREN: A Sustainable Cache for Emerging Write-Read-Erase Flash InterfacesSara McAllister, Yucong Wang, Benjamin Berg, Daniel S. Berger 等OSDI 2024 · 被引用 14 次
- Siloz: Leveraging DRAM Isolation Domains to Prevent Inter-VM RowhammerKevin Loughlin, Jonah Rosenblum, Stefan Saroiu, Alec Wolman 等SOSP 2023 · 被引用 13 次
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock 等ISCA 2024 · 被引用 12 次
它引用的顶会 Paper8
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun 等OSDI 2020 · 被引用 101 次
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 被引用 69 次
相关 Paper
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu 等HPCA 2021 · 被引用 9 次
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry 等ASPLOS 2022 · 被引用 39 次
- Flex: High-Availability Datacenters With Zero Reserved PowerChaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang 等ISCA 2021 · 被引用 43 次
- Fault Escaping: Improving Robustness of DPU Enhanced Platform with Mutual Assisted VM RecoveryChao Zhang, Tao Xu, Junming Liu, Pai Liu 等ASPLOS 2025
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte 等ASPLOS 2020 · 被引用 58 次
