USENIX ATC2021顶会
Fighting the Fog of War: Automated Incident Detection for Cloud Systems
Liqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang, Yu Kang, Pu Zhao, Bo Qiao, Shilin He, Pochian Lee, Jeffrey Sun, Feng Gao, Li Yang
摘要
Incidents and outages dramatically degrade the availability of large-scale cloud computing systems such as AWS, Azure, and GCP. In current incident response practice, each team has only a partial view of the entire system, which makes the detection of incidents like fighting in the "fog of war". As a result, prolonged mitigation time and more financial loss are incurred. In this work, we propose an automatic incident detection system, namely Warden, as a part of the Incident Management (IcM) platform. Warden collects alerts from different services and detects the occurrence of incidents from a global perspective. For each detected potential incident, Warden notifies relevant on-call engineers so that they could properly prioritize their tasks and initiate cross-team collaboration. We implemented and deployed Warden in the IcM platform of Azure. Our evaluation results based on data collected in an 18-month period from 26 major services show that Warden is effective and outperforms the baseline methods. For the majority of successfully detected incidents ( 68%), Warden is faster than human, and this is particularly the case for the incidents that take long time to detect manually.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Xpert: Empowering Incident Management with Query Recommendations via Large Language ModelsYuxuan Jiang, Chaoyun Zhang, Shilin He, Zhihao Yang 等ICSE 2024 · 被引用 24 次
- Incident Response Planning Using a Lightweight Large Language Model with Reduced HallucinationKim Hammar, Tansu Alpcan, Emil C. LupuNDSS 2026 · 被引用 16 次
- Incident-aware Duplicate Ticket Aggregation for Cloud SystemsJinyang Liu, Shilin He, Zhuangbin Chen, Liqun Li 等ICSE 2023 · 被引用 14 次
- Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud SystemsJinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang 等ASE 2023 · 被引用 7 次
- Outage-Watch: Early Prediction of Outages using Extreme Event RegularizerShubham Agarwal, Sarthak Chakraborty, Shaddy Garg, Sumit Bisht 等FSE 2023 · 被引用 5 次
它引用的顶会 Paper4
- AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyMeng Ma, Jingmin Xu, Yuan Wang, Pengfei Chen 等WWW 2020 · 被引用 144 次
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He 等FSE 2020 · 被引用 43 次
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin 等ASE 2020 · 被引用 33 次
相关 Paper
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang 等ICSE 2021 · 被引用 24 次
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia 等ASE 2025 · 被引用 2 次
- Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchJiazhen Gu, Chuan Luo, Si Qin, Bo Qiao 等FSE 2020 · 被引用 35 次
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann 等ICSE 2023 · 被引用 93 次
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
