AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou, Kun Wang, Shuicheng Yan
Abstract
Large Language Model (LLM)-based agentic systems, often comprising multiple models, complex tool invocations, and orchestration protocols, substantially outperform monolithic agents. Yet this very sophistication amplifies their fragility, making them more prone to system failure. Pinpointing the specific agent or step responsible for an error within long execution traces defines the task of agentic system failure attribution. Current state-of-the-art reasoning LLMs, however, remain strikingly inadequate for this challenge, with accuracy generally below 10%. To address this gap, we propose AgenTracer, the first automated framework for annotating failed multi-agent trajectories via counterfactual replay and programmed fault injection, producing the curated dataset TracerTraj. Leveraging this resource, we develop AgenTracer-8B, a lightweight failure tracer trained with multi-granular reinforcement learning, capable of efficiently diagnosing errors in verbose multi-agent interactions. On Who&When benchmark, AgenTracer-8B outperforms giant proprietary LLMs like Gemini-2.5-Pro and Claude-4-Sonnet by up 18.18%, setting a new standard in LLM agentic failure attribution. More importantly, AgenTracer-8B delivers actionable feedback to off-the-shelf multi-agent systems like MetaGPT and MaAS with 4.8 ∼ 14.2% performance gains, empowering self-correcting and self-evolving agentic AI. Our project page is at https://bingreeky.github.io/atracer/ . Who&When (handcraft; agent-level) Who&When (handcraft; step-level) TracerTraj-Code (agent-level) TracerTraj-Math (step-level)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c670c58-321c-4b4f-b265-0c9c730efe18Cited by top-tier papers4
- DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyYixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun et al.ICLR 2026 · 57 citations
- Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent SystemsFulin Lin, Shaowen Chen, Ruishan Fang, Hongwei Wang et al.ICLR 2026 · 15 citations
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge GraphsYurun Chen, Xueyu Hu, Yuhan Liu, Ziqi Wang et al.CVPR 2026 · 7 citations
- StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent SystemsTaiyu Zhu, Yifan Wu, Weilin Jin, Ying Li et al.KDD 2026 · 2 citations
Builds on23
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
Related papers
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent SystemsShaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu et al.ICML 2025
- Aegis: Automated Error Generation and Attribution for Multi-Agent SystemsFanqi Kong, Ruijie Zhang, Huaxiao Yin, Guibin Zhang et al.ICLR 2026 · 16 citations
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent SystemsMengzhuo Chen, Junjie Wang, Fangwen Mu, Yawen Wang et al.ACL 2026 · 5 citations
- Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent SystemsKai Sun, Wenqiang Li, Bo Dong, Yuxin Lin et al.AAAI 2026
- Spectrum-Based Failure Attribution for Multi-agent SystemsYu Ge, Linna Xie, Zhong Li, Yu Pei et al.FSE 2026
