Fast Outage Analysis of Large-scale Production Clouds with Service Correlation Mining
Yaohui Wang, Guozheng Li, Zijian Wang, Yu Kang, Yangfan Zhou, Hongyu Zhang, Feng Gao, Jeffrey Sun, Li Yang, Pochian Lee, Zhangwei Xu, Pu Zhao
Abstract
Cloud-based services are surging into popularity in recent years. However, outages, i.e., severe incidents that always impact multiple services, can dramatically affect user experience and incur severe economic losses. Locating the root-cause service, i.e., the service that contains the root cause of the outage, is a crucial step to mitigate the impact of the outage. In current industrial practice, this is generally performed in a bootstrap manner and largely depends on human efforts: the service that directly causes the outage is identified first, and the suspected root cause is traced back manually from service to service during diagnosis until the actual root cause is found. Unfortunately, production cloud systems typically contain a large number of interdependent services. Such a manual root cause analysis is often time-consuming and labor-intensive. In this work, we propose COT, the first outage triage approach that considers the global view of service correlations. COT mines the correlations among services from outage diagnosis data. After learning from historical outages, COT can infer the root cause of emerging ones accurately. We implement COT and evaluate it on a real-world dataset containing one year of data collected from Microsoft Azure, one of the representative cloud computing platforms in the world. Our experimental results show that COT can reach a triage accuracy of 82.1%-83.5%, which outperforms the state-of-the-art triage approach by 28.0%-29.7%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cad226f-3708-477f-8f74-d0ea6426a7cbCited by top-tier papers5
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling et al.ASE 2021 · 23 citations
- Incident-aware Duplicate Ticket Aggregation for Cloud SystemsJinyang Liu, Shilin He, Zhuangbin Chen, Liqun Li et al.ICSE 2023 · 14 citations
- Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud SystemsJinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang et al.ASE 2023 · 7 citations
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen et al.ICSE 2025 · 4 citations
- ESRO: Experience Assisted Service Reliability against OutagesSarthak Chakraborty, Shubham Agarwal, Shaddy Garg, Abhimanyu Sethia et al.ASE 2023 · 3 citations
Related papers
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang et al.USENIX ATC 2021 · 43 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia et al.ASE 2025 · 2 citations
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- AutoARTS: Taxonomy, Insights and Tools for Root Cause Labelling of Incidents in Microsoft AzurePradeep Dogga, Chetan Bansal, Richard Costleigh, Gopinath Jayagopal et al.USENIX ATC 2023 · 14 citations
