FoundRoot: Towards Foundation Model for Root Cause Analysis via Structured Deep Thinking
Zhe Xie, Zeyan Li, Xiao He, Shenglin Zhang, Longlong Xu, Yuzhuo Yang, Tieying Zhang, Jianjun Chen, Rui Shi, Dan Pei
摘要
Root Cause Analysis (RCA) for service systems is critical for ensuring their reliability, while its application remains challenging because of the large number of metrics and the complex causal relationships. Classical RCA methods typically rely on statistical or rule-based approaches, making them difficult to generalize to unseen systems. The introduction of Large Language Models (LLMs) has partly addressed these challenges with their understanding of the domain-specific semantics of metrics and reasoning capabilities. However, they still struggle with incomplete or shallow reasoning when facing a large amount of metrics. To address these limitations, we present FoundRoot, a reinforcement learning (RL)-enhanced LLM foundation model for zero-shot RCA. FoundRoot features a novel structured deep thinking paradigm, which breaks down the RCA reasoning into several goal-oriented substeps, enhancing the completeness and reasoning capability of the causal relationship. We collect and curate diverse open-sourced RCA datasets across different systems and introduce a data augmentation technique to ensure data scalability. We design a two-stage training pipeline that includes supervised fine-tuning (SFT) and RL to align the LLM with the structured deep thinking paradigm, which significantly improves its reasoning quality. Extensive experiments on four datasets show that FoundRoot outperforms both classical and LLM-based methods in RCA accuracy on unseen systems, achieving 4.5%-48.6% mean reciprocal rank (MRR) improvements. The source code and data of this paper is available at: https://github.com/NetManAIOps/FoundRoot.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini 等NeurIPS 2022 · 被引用 185 次
- Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability DataGuangba Yu, Pengfei Chen, Yufeng Li, Hongyang Chen 等FSE 2023 · 被引用 131 次
- Eadro: An End-to-End Troubleshooting Framework for Microservices on Multi-source DataCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su 等ICSE 2023 · 被引用 99 次
相关 Paper
- RECoRD: A Multi-Agent LLM Framework for Reverse Engineering Codebase to Relational DiagramYuan Xue, Xiaoyu Lu, Yunfei Bai, Yunan Liu 等AAAI 2026
- The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small ClassifierYongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu 等ASE 2024 · 被引用 3 次
- OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?Junjielong Xu, Qinan Zhang, Zhiqing Zhong, Shilin He 等ICLR 2025
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal KnowledgeShuai Liang, Pengfei Chen, Bozhe Tian, Gou Tan 等FSE 2026 · 被引用 4 次
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang 等NeurIPS 2025 · 被引用 153 次
