FoundRoot: Towards Foundation Model for Root Cause Analysis via Structured Deep Thinking
Zhe Xie, Zeyan Li, Xiao He, Shenglin Zhang, Longlong Xu, Yuzhuo Yang, Tieying Zhang, Jianjun Chen, Rui Shi, Dan Pei
Abstract
Root Cause Analysis (RCA) for service systems is critical for ensuring their reliability, while its application remains challenging because of the large number of metrics and the complex causal relationships. Classical RCA methods typically rely on statistical or rule-based approaches, making them difficult to generalize to unseen systems. The introduction of Large Language Models (LLMs) has partly addressed these challenges with their understanding of the domain-specific semantics of metrics and reasoning capabilities. However, they still struggle with incomplete or shallow reasoning when facing a large amount of metrics. To address these limitations, we present FoundRoot, a reinforcement learning (RL)-enhanced LLM foundation model for zero-shot RCA. FoundRoot features a novel structured deep thinking paradigm, which breaks down the RCA reasoning into several goal-oriented substeps, enhancing the completeness and reasoning capability of the causal relationship. We collect and curate diverse open-sourced RCA datasets across different systems and introduce a data augmentation technique to ensure data scalability. We design a two-stage training pipeline that includes supervised fine-tuning (SFT) and RL to align the LLM with the structured deep thinking paradigm, which significantly improves its reasoning quality. Extensive experiments on four datasets show that FoundRoot outperforms both classical and LLM-based methods in RCA accuracy on unseen systems, achieving 4.5%-48.6% mean reciprocal rank (MRR) improvements. The source code and data of this paper is available at: https://github.com/NetManAIOps/FoundRoot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 603b31cb-6ddd-439c-b4f4-33b4c276daffBuilds on16
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
- Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability DataGuangba Yu, Pengfei Chen, Yufeng Li, Hongyang Chen et al.FSE 2023 · 131 citations
- Eadro: An End-to-End Troubleshooting Framework for Microservices on Multi-source DataCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su et al.ICSE 2023 · 99 citations
Related papers
- RECoRD: A Multi-Agent LLM Framework for Reverse Engineering Codebase to Relational DiagramYuan Xue, Xiaoyu Lu, Yunfei Bai, Yunan Liu et al.AAAI 2026
- The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small ClassifierYongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu et al.ASE 2024 · 3 citations
- OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?Junjielong Xu, Qinan Zhang, Zhiqing Zhong, Shilin He et al.ICLR 2025
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal KnowledgeShuai Liang, Pengfei Chen, Bozhe Tian, Gou Tan et al.FSE 2026 · 4 citations
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
