AutoARTS: Taxonomy, Insights and Tools for Root Cause Labelling of Incidents in Microsoft Azure
Pradeep Dogga, Chetan Bansal, Richard Costleigh, Gopinath Jayagopal, Suman Nath, Xuchao Zhang
Abstract
Labelling incident postmortems with the root causes is essential for aggregate analysis, which can reveal common problem areas, trends, patterns, and risks that may cause future incidents. A common practice is to manually label postmortems with a single root cause based on an ad hoc taxonomy of root cause tags. However, this manual process is error-prone, a single root cause is inadequate to capture all contributing factors behind an incident, and ad hoc taxonomies do not reflect the diverse categories of root causes.
In this paper, we address this problem with a three-pronged approach. First, we conduct an extensive multi-year analysis of over 2000 incidents from more than 450 services in Microsoft Azure to understand all the factors that contributed to the incidents. Second, based on the empirical study, we propose a novel hierarchical and comprehensive taxonomy of potential contributing factors for production incidents. Lastly, we develop an automated tool that can assist humans in the labelling process. We present empirical evaluation and a user study that show the effectiveness of our approach. To the best of our knowledge, this is the largest and most comprehensive study of production incident postmortem reports yet. We also make our taxonomy publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a00abd1-5131-403f-9507-4336cd91e951Cited by top-tier papers1
Ask how each one uses itBuilds on7
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
- Hierarchy-Aware Global Model for Hierarchical Text ClassificationJie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu et al.ACL 2020 · 171 citations
- Sequence Level Contrastive Learning for Text SummarizationShusheng Xu, Xingxing Zhang, Yi Wu, Furu WeiAAAI 2022 · 113 citations
- Rex: Preventing Bugs and Misconfiguration in Large Services Using Correlated Change AnalysisSonu Mehta, Ranjita Bhagwan, Rahul Kumar, Chetan Bansal et al.NSDI 2020 · 65 citations
Related papers
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang et al.ICSE 2021 · 24 citations
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchJiazhen Gu, Chuan Luo, Si Qin, Bo Qiao et al.FSE 2020 · 35 citations
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin et al.ASE 2020 · 33 citations
