Hyrax: Fail-in-Place Server Operation in Cloud Platforms
Jialun Lyu, Marisa You, Celine Irvene, Mark Jung, Tyler Narmore, Jacob Shapiro, Luke Marshall, Savyasachi Samal, Ioannis Manousakis, Lisa Hsu, Preetha Subbarayalu, Ashish Raniwala
Abstract
Today's cloud platforms handle server hardware failures by shutting down the affected server and only turning it back online once it has been repaired by a technician. At cloud scale, this all-or-nothing operating model is becoming increasingly unsustainable. This model is also at odds with technology trends, such as the need for new cooling technology.
This paper introduces Hyrax, a datacenter stack that enables compute servers with failed components to continue hosting VMs while hiding the underlying degraded capacity and performance. A key enabler of Hyrax is a novel model of changes in memory interleaving when deactivating faulty memory modules. Experiments on cloud production servers show that Hyrax overcomes common hardware failures without impacting peak VM performance. In large-scale simulations with production traces, Hyrax reduces server repair requirements by 50-60% without impacting VM scheduling.
- Formerly at Microsoft Azure 1 Prior work reported 0.09% [58, 59], 0.12% [12], and 1.6% [57,66] for DIMMs and 0.22% [43,44] to 1.2% [3,54] for SSDs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5da54c78-c902-4a00-8283-e5e91ae8a3fdCited by top-tier papers8
- Understanding Silent Data Corruptions in a Large Production CPU PopulationShaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang et al.SOSP 2023 · 56 citations
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene et al.ISCA 2024 · 49 citations
- FairyWREN: A Sustainable Cache for Emerging Write-Read-Erase Flash InterfacesSara McAllister, Yucong Wang, Benjamin Berg, Daniel S. Berger et al.OSDI 2024 · 14 citations
- Siloz: Leveraging DRAM Isolation Domains to Prevent Inter-VM RowhammerKevin Loughlin, Jonah Rosenblum, Stefan Saroiu, Alec Wolman et al.SOSP 2023 · 13 citations
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock et al.ISCA 2024 · 12 citations
Builds on8
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun et al.OSDI 2020 · 101 citations
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 69 citations
Related papers
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu et al.HPCA 2021 · 9 citations
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry et al.ASPLOS 2022 · 39 citations
- Flex: High-Availability Datacenters With Zero Reserved PowerChaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang et al.ISCA 2021 · 43 citations
- Fault Escaping: Improving Robustness of DPU Enhanced Platform with Mutual Assisted VM RecoveryChao Zhang, Tao Xu, Junming Liu, Pai Liu et al.ASPLOS 2025
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte et al.ASPLOS 2020 · 58 citations
