Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, Qingyun Wu
摘要
Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we propose and formulate a new research area: automated failure attribution for LLM multi-agent systems. To support this initiative, we introduce the Who&When dataset, comprising extensive failure logs from 127 LLM multi-agent systems with fine-grained annotations linking failures to specific agents and decisive error steps. Using the Who&When, we develop and evaluate three automated failure attribution methods, summarizing their corresponding pros and cons. The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps, with some methods performing below random. Even SOTA reasoning models, such as OpenAI o1 and DeepSeek R1, fail to achieve practical usability. These results highlight the task's complexity and the need for further research in this area. Code and dataset are available in the public repository.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou 等ICLR 2026 · 被引用 107 次
- DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyYixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun 等ICLR 2026 · 被引用 57 次
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su 等NeurIPS 2025 · 被引用 35 次
- CoAct-1: Computer-using Multi-agent System with Coding ActionsLinxin Song, Yutong Dai, Viraj Prabhu, Jieyu Zhang 等ICLR 2026 · 被引用 32 次
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang 等ICLR 2026 · 被引用 24 次
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
相关 Paper
- Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent SystemsKai Sun, Wenqiang Li, Bo Dong, Yuxin Lin 等AAAI 2026
- StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent SystemsTaiyu Zhu, Yifan Wu, Weilin Jin, Ying Li 等KDD 2026 · 被引用 2 次
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent SystemsMengzhuo Chen, Junjie Wang, Fangwen Mu, Yawen Wang 等ACL 2026 · 被引用 5 次
- Spectrum-Based Failure Attribution for Multi-agent SystemsYu Ge, Linna Xie, Zhong Li, Yu Pei 等FSE 2026
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin 等CHI 2026 · 被引用 2 次
