Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure
Ze Li, Qian Cheng, Ken Hsieh, Yingnong Dang, Peng Huang, Pankaj Singh, Xinsheng Yang, Qingwei Lin, Youjiang Wu, Sebastien Levy, Murali Chintalapati
摘要
Modern cloud systems have a vast number of components that continuously undergo updates. Deploying these frequent updates quickly without breaking the system is challenging. In this paper, we present Gandalf, an end-to-end analytics service for safe deployment in a large-scale system infrastructure. Gandalf enables rapid and robust impact assessment of software rollouts to catch bad rollouts before they cause widespread outages. Gandalf monitors and analyzes various fault signals. It will correlate each signal against all the ongoing rollouts using a spatial and temporal correlation algorithm. The core decision logic of Gandalf includes an ensemble ranking algorithm that determines which rollout may have caused the fault signals, and a binary classifier that assesses the impact of the fault signals. The analysis result will decide whether a rollout is safe to proceed or should be stopped.
By using a lambda architecture, Gandalf provides both realtime and long-term deployment monitoring with automated decisions and notifications. Gandalf has been running in production in Microsoft Azure for more than 18 months, serving both data-plane and control-plane components. It achieves 92.4% precision and 100% recall (no high-impact service outages in Azure Compute were caused by bad rollouts) for dataplane rollouts. For control-plane rollouts, Gandalf achieves 94.9% precision and 99.8% recall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang 等FSE 2021 · 被引用 89 次
- Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal AttentionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su 等ICSE 2023 · 被引用 52 次
- Understanding and Detecting Software Upgrade Failures in Distributed SystemsYongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi 等SOSP 2021 · 被引用 40 次
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu 等WWW 2021 · 被引用 37 次
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang 等OSDI 2020 · 被引用 35 次
它引用的顶会 Paper2
相关 Paper
- Check before You Change: Preventing Correlated Failures in Service UpdatesEnnan Zhai, Ang Chen, Ruzica Piskac, Mahesh Balakrishnan 等NSDI 2020 · 被引用 46 次
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling 等ASE 2021 · 被引用 23 次
- Aragog: Scalable Runtime Verification of Shardable Networked SystemsNofel Yaseen, Behnaz Arzani, Ryan Beckett, Selim Ciraci 等OSDI 2020 · 被引用 19 次
- Fidelity of Cloud Emulators: The Imitation Game of Testing Cloud-Based SoftwareAnna Mazhar, Saad Sher Alam, William X. Zheng, Yinfang Chen 等ICSE 2025 · 被引用 2 次
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang 等ICSE 2021 · 被引用 24 次
