Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions
Sebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang, Peng Huang, Zheng Mu, Pu Zhao, Tarun Ramani, Naga K. Govindaraju, Xukun Li, Qingwei Lin, Gil Lapid Shafriri, Murali Chintalapati
Abstract
This technical report is an extended version of our OSDI 2020 paper: Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions. Abstract: When a failure occurs in production systems, the highest priority is to quickly mitigate it. Despite its importance, failure mitigation is done in a reactive and ad-hoc way: taking some fixed actions only after a severe symptom is observed. For cloud systems, such a strategy is inadequate. In this paper, we propose a preventive and adaptive failure mitigation service, NARYA, that is integrated in a production cloud, Microsoft Azure’s compute platform. Narya predicts imminent host failures based on multi-layer system signals and then decides smart mitigation actions. The goal is to avert VM failures. Narya’s decision engine takes a novel online experimentation approach to continually explore the best mitigation action. Narya further enhances the adaptive decision capability through reinforcement learning. Narya has been running in production for 15 months. It on average reduces VM interruptions by 26% compared to previous static strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e4d2b2b-6dec-4f16-a7f8-15a9ea9665f8Cited by top-tier papers9
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 220 citations
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang et al.FSE 2021 · 89 citations
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu et al.WWW 2021 · 37 citations
- RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud InfrastructureChang Lou, Cong Chen, Peng Huang, Yingnong Dang et al.OSDI 2022 · 18 citations
- Demystifying and Checking Silent Semantic Violations in Large Distributed SystemsChang Lou, Yuzhuo Jing, Peng HuangOSDI 2022 · 8 citations
Builds on4
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao et al.FAST 2020 · 120 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud InfrastructureZe Li, Qian Cheng, Ken Hsieh, Yingnong Dang et al.NSDI 2020 · 69 citations
- Meaningful AvailabilityTamas Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean et al.NSDI 2020 · 20 citations
Related papers
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang et al.USENIX ATC 2021 · 43 citations
- Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the FieldIsaac Boixaderas, Sergi Moré, Javier Bartolome, David Vicente et al.HPDC 2024 · 1 citation
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su et al.NeurIPS 2025 · 35 citations
- SOL: safe on-node learning in cloud platformsYawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger et al.ASPLOS 2022 · 14 citations
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang et al.ICSE 2021 · 24 citations
