RCAFlow: A Workflow-Informed Hierarchical Planning Multi-Agent System for Root Cause Analysis
Yufei Gao, Zhengong Cai, Bowei Yang
摘要
As microservice architectures become increasingly complex and system events become more frequent, Root Cause Analysis (RCA) has emerged as a critical task to ensure system reliability. However, existing deep learning-based methods often struggle with limited flexibility and a lack of interpretability when addressing complex system failures. Recent efforts to integrate large language models (LLMs) have shown promise in enhancing diagnostic transparency and reasoning capability. However, expansive search spaces, intricate workflows, and entangled constraints constrain practical adoption. We propose RCAFlow, a multi-agent framework that integrates structured workflow knowledge with hierarchical planning to address these challenges. RCAFlow transforms semistructured documents into behavior tree-style workflows to support interpretable plan generation, employs a Git-inspired branching mechanism for modular and hierarchical task execution with path isolation, and leverages state-aware task execution with semantic analysis to improve result understanding and feedback. We evaluate RCAFlow on three benchmark datasets provided by OpenRCA. Experimental results demonstrate that RCAFlow consistently outperforms existing methods across all datasets. Further ablation studies confirm the effectiveness of each core module, highlighting the reliability, extensibility, and interpretability of RCAFlow to support complex RCA tasks within intelligent IT operations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Anomaly Transformer: Time Series Anomaly Detection with Association DiscrepancyJiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng LongICLR 2022 · 被引用 960 次
- AdaPlanner: Adaptive Planning from Feedback with Language ModelsHaotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai 等NeurIPS 2023 · 被引用 257 次
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann 等ICSE 2023 · 被引用 93 次
相关 Paper
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal KnowledgeShuai Liang, Pengfei Chen, Bozhe Tian, Gou Tan 等FSE 2026 · 被引用 4 次
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin 等CHI 2026 · 被引用 2 次
- RECoRD: A Multi-Agent LLM Framework for Reverse Engineering Codebase to Relational DiagramYuan Xue, Xiaoyu Lu, Yunfei Bai, Yunan Liu 等AAAI 2026
- COCA: Generative Root Cause Analysis for Distributed Systems with Code KnowledgeYichen Li, Yulun Wu, Jinyang Liu, Zhihan Jiang 等ICSE 2025 · 被引用 6 次
- The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small ClassifierYongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu 等ASE 2024 · 被引用 3 次
