Lune

ISSTA2026Top-tier venue

AgentInspect: Diagnosing Behavioral Failures in Artificial Intelligence Agents

Ruchira Manke, Mohammad Wardat, Foutse Khomh, Hridesh Rajan

2026Year

Abstract

Effectively testing Artificial Intelligence (AI) agents remains a fundamental challenge due to their stochastic reasoning, vast and diverse input space, reliance on external tools, and operation in dynamic execution environments; factors that demand new testing methodologies explicitly tailored to the complex and interactive nature of agent-based systems. This work presents a novel methodology for testing AI agents, with a particular focus on assessing their behavioral robustness under varied operational conditions. Our approach relies on following key technical innovations: (1) a coverage-guided test input generation strategy based on agent- specific coverage objectives, (2) a capture-and-simulate mechanism that systematically emulates abnormal tool behaviors to mimic real-world execution failures, and (3) a deterministic behavioral failure detection approach that enables consistent identification of failures across different test inputs. We developed AgentInspect, a framework that automatically detects six types of behavioral failures in LangChain-based AI agents by analyzing their execution trajectories across three evaluation settings: a baseline setting using real tool responses, a simulated setting incorporating synthetic tool responses, and a hybrid setting that combines the real and simulated tool responses. To evaluate our approach, we curated a benchmark of 35 AI agents obtained from GitHub. Our results show that AgentInspect consistently identifies different behavioral failures with high precision and recall across all three execution settings. In particular, the simulated and hybrid settings expose failure modes that do not emerge during baseline execution with real tool responses, thereby enabling a more comprehensive assessment of agent robustness. Our findings highlight AgentInspect’s effectiveness in revealing critical failures and its practical utility for systematic robustness evaluation of AI agents.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 7b2f1fa1-5917-4d29-acc2-e36ba5766643

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines