AgentInspect: Diagnosing Behavioral Failures in Artificial Intelligence Agents
Ruchira Manke, Mohammad Wardat, Foutse Khomh, Hridesh Rajan
摘要
Effectively testing Artificial Intelligence (AI) agents remains a fundamental challenge due to their stochastic reasoning, vast and diverse input space, reliance on external tools, and operation in dynamic execution environments; factors that demand new testing methodologies explicitly tailored to the complex and interactive nature of agent-based systems. This work presents a novel methodology for testing AI agents, with a particular focus on assessing their behavioral robustness under varied operational conditions. Our approach relies on following key technical innovations: (1) a coverage-guided test input generation strategy based on agent- specific coverage objectives, (2) a capture-and-simulate mechanism that systematically emulates abnormal tool behaviors to mimic real-world execution failures, and (3) a deterministic behavioral failure detection approach that enables consistent identification of failures across different test inputs. We developed AgentInspect, a framework that automatically detects six types of behavioral failures in LangChain-based AI agents by analyzing their execution trajectories across three evaluation settings: a baseline setting using real tool responses, a simulated setting incorporating synthetic tool responses, and a hybrid setting that combines the real and simulated tool responses. To evaluate our approach, we curated a benchmark of 35 AI agents obtained from GitHub. Our results show that AgentInspect consistently identifies different behavioral failures with high precision and recall across all three execution settings. In particular, the simulated and hybrid settings expose failure modes that do not emerge during baseline execution with real tool responses, thereby enabling a more comprehensive assessment of agent robustness. Our findings highlight AgentInspect’s effectiveness in revealing critical failures and its practical utility for systematic robustness evaluation of AI agents.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LogicHunter: Testing LLM Agent Frameworks with an Agentic OracleMinghui Long, Yanjie Zhao, Haoyu WangISSTA 2026
- SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI EnvironmentsSyed Yusuf Ahmed, Shiwei Feng, Chanwoo Bae, Calix Barrus 等ICSE 2026
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren 等ICML 2026 · 被引用 5 次
- A Computational Framework for Evaluating Human-likeness in LLMs' Open-ended Human BehaviorsYuxuan Lei, Jianxun Lian, Defu Lian, Jincenzi Wu 等ICML 2026
- AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment CorruptionsJingwei Sun, Jianing Zhu, Yuanyi Li, Tongliang Liu 等ICML 2026 · 被引用 2 次
