Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions
Sebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang, Peng Huang, Zheng Mu, Pu Zhao, Tarun Ramani, Naga K. Govindaraju, Xukun Li, Qingwei Lin, Gil Lapid Shafriri, Murali Chintalapati
摘要
This technical report is an extended version of our OSDI 2020 paper: Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions. Abstract: When a failure occurs in production systems, the highest priority is to quickly mitigate it. Despite its importance, failure mitigation is done in a reactive and ad-hoc way: taking some fixed actions only after a severe symptom is observed. For cloud systems, such a strategy is inadequate. In this paper, we propose a preventive and adaptive failure mitigation service, NARYA, that is integrated in a production cloud, Microsoft Azure’s compute platform. Narya predicts imminent host failures based on multi-layer system signals and then decides smart mitigation actions. The goal is to avert VM failures. Narya’s decision engine takes a novel online experimentation approach to continually explore the best mitigation action. Narya further enhances the adaptive decision capability through reinforcement learning. Narya has been running in production for 15 months. It on average reduces VM interruptions by 26% compared to previous static strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 被引用 220 次
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang 等FSE 2021 · 被引用 89 次
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu 等WWW 2021 · 被引用 37 次
- RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud InfrastructureChang Lou, Cong Chen, Peng Huang, Yingnong Dang 等OSDI 2022 · 被引用 18 次
- Demystifying and Checking Silent Semantic Violations in Large Distributed SystemsChang Lou, Yuzhuo Jing, Peng HuangOSDI 2022 · 被引用 8 次
它引用的顶会 Paper4
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao 等FAST 2020 · 被引用 120 次
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud InfrastructureZe Li, Qian Cheng, Ken Hsieh, Yingnong Dang 等NSDI 2020 · 被引用 69 次
- Meaningful AvailabilityTamas Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean 等NSDI 2020 · 被引用 20 次
相关 Paper
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang 等USENIX ATC 2021 · 被引用 43 次
- Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the FieldIsaac Boixaderas, Sergi Moré, Javier Bartolome, David Vicente 等HPDC 2024 · 被引用 1 次
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su 等NeurIPS 2025 · 被引用 35 次
- SOL: safe on-node learning in cloud platformsYawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger 等ASPLOS 2022 · 被引用 14 次
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang 等ICSE 2021 · 被引用 24 次
