Automatic Root Cause Analysis via Large Language Models for Cloud Incidents
Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang, Xin Gao, Liu Shi, Yunjie Cao, Xuedong Gao, Hao Fan, Ming Wen, Jun Zeng, Supriyo Ghosh
Abstract
Ensuring the reliability and availability of cloud services necessitates efficient root cause analysis (RCA) for cloud incidents. Traditional RCA methods, which rely on manual investigations of data sources such as logs and traces, are often laborious, error-prone, and challenging for on-call engineers. In this paper, we introduce RCACopilot, an innovative on-call system empowered by the large language model for automating RCA of cloud incidents. RCACopilot matches incoming incidents to corresponding incident handlers based on their alert types, aggregates the critical runtime diagnostic information, predicts the incident's root cause category, and provides an explanatory narrative. We evaluate RCACopilot using a real-world dataset consisting of a year's worth of incidents from Microsoft. Our evaluation demonstrates that RCACopilot achieves RCA accuracy up to 0.766. Furthermore, the diagnostic information collection component of RCACopilot has been successfully in use at Microsoft for over four years.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43fae99b-e217-4057-b51d-fe477d534b79Cited by top-tier papers31
- Revisiting VAE for Unsupervised Time Series Anomaly Detection: A Frequency PerspectiveZexin Wang, Changhua Pei, Minghua Ma, Xin Wang et al.WWW 2024 · 90 citations
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su et al.NeurIPS 2025 · 35 citations
- NetArena: Dynamic Benchmarks for AI Agents in Network AutomationYajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi et al.ICLR 2026 · 17 citations
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin et al.SIGCOMM 2025 · 13 citations
- ART: A Unified Unsupervised Framework for Incident Management in Microservice SystemsYongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma et al.ASE 2024 · 9 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen et al.FSE 2022 · 206 citations
- UniParser: A Unified Log Parser for Heterogeneous Log DataYudong Liu, Xu Zhang, Shilin He, Hongyu Zhang et al.WWW 2022 · 148 citations
Related papers
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?Junjielong Xu, Qinan Zhang, Zhiqing Zhong, Shilin He et al.ICLR 2025
- LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call AutomationRuowei Fu, Yang Zhang, Zeyu Che, Xin Wu et al.ASE 2025
- The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small ClassifierYongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu et al.ASE 2024 · 3 citations
- COCA: Generative Root Cause Analysis for Distributed Systems with Code KnowledgeYichen Li, Yulun Wu, Jinyang Liu, Zhihan Jiang et al.ICSE 2025 · 6 citations
