SkillTracer: Structural Failure Attribution and Refinement of Agentic Skills in Long-Horizon Web Tasks
Yuyang Li, Yiran Dou, Jie-Jing Shao, Yueming Lyu, Ivor Tsang, Haiyan Yin
摘要
Long-horizon web agents frequently fail without knowing where or why execution broke down. This issue is particularly pronounced in skill-based agentic web systems, where failures arise within composite skills whose internal decision processes are not directly traceable, making precise diagnosis and repair especially difficult over long horizons. We introduce SkillTracer, a framework that represents skills as attributed plan graphs structured by hierarchical nodes and verifiable edge transitions, enabling programmatic verification of execution progress. By decomposing skills into inspectable hierarchies, SkillTracer converts raw interaction traces into structural evidence, making execution breakdowns localizable to specific node-level decision points and attributable to failing components. This attribution signal facilitates targeted structural repair, allowing the agent to selectively revise failing components while preserving the integrity of valid substructures for partial reuse and adaptive recovery. Furthermore, SkillTracer synthesizes short-term traces with long-term historical evidence to construct a persistent skill graph, enabling failure patterns to drive continual refinement across episodes. Evaluated on challenging long-horizon benchmarks, SkillTracer achieves a 17.7% average improvement in success rate over strong baselines, with gains of up to 56.3% in cross-domain settings, demonstrating that structural attribution and skill repair are critical for reliable long-horizon web interaction. A project page is available at: https://liyuuuuy.github.io/SkillTracer/.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning PerspectiveMohamed Aghzal, Gregory J. Stein, Ziyu YaoACL 2026 · 被引用 5 次
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou 等ICLR 2026 · 被引用 107 次
- GraphSkill: Documentation-Guided Agentic Hierarchical Retrieval-Augmented Coding for Complex Graph ReasoningFali Wang, Chenglin Weng, Xianren Zhang, Siyuan Hong 等KDD 2026
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent SystemsMengzhuo Chen, Junjie Wang, Fangwen Mu, Yawen Wang 等ACL 2026 · 被引用 5 次
- Creating Multi-Level Skill Hierarchies in Reinforcement LearningJoshua B. Evans, Özgür SimsekNeurIPS 2023 · 被引用 15 次
