LogicHunter: Testing LLM Agent Frameworks with an Agentic Oracle
Minghui Long, Yanjie Zhao, Haoyu Wang
摘要
Large Language Model (LLM) agent frameworks such as LangChain, LlamaIndex, and CrewAI have become critical infrastructure powering production AI systems, yet they remain severely under-tested due to fundamental challenges in automated testing. Unlike traditional software, where crashes serve as reliable oracles, defects in these pure Python frameworks manifest as ordinary exceptions or silent semantic failures, creating profound oracle ambiguity. This problem is exacerbated by strict type governance through Pydantic schemas and complex protocol requirements that cause existing fuzzers to generate overwhelming invalid inputs, while traditional test generators produce only trivial cases with weak regression assertions.
We present LogicHunter, a fuzzing framework that addresses both the generation and oracle challenges through active specification-aware testing. LogicHunter employs specification-driven generation that systematically fuses formal type constraints with authentic usage patterns from real-world repositories, synthesizing inputs that are valid by construction yet semantically extreme, equipped with behavioral probes to expose silent failures. To resolve oracle ambiguity, we introduce the Agentic Oracle, which transcends passive classification by actively retrieving documentation, navigating source code, and inspecting runtime states through a ReAct-based architecture with Dual-Layer State Management and Dual-Stream Memory. Evaluated on three widely deployed frameworks, LogicHunter discovered 40 previously unknown bugs with 30 confirmed and 26 fixed by developers, while state-of-the-art baselines reported no bugs as final findings. The Agentic Oracle achieves 91.17% precision, surpassing the best passive approach at 29.27% by 61 percentage points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang 等ISSTA 2023 · 被引用 253 次
相关 Paper
- SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI EnvironmentsSyed Yusuf Ahmed, Shiwei Feng, Chanwoo Bae, Calix Barrus 等ICSE 2026
- Make Agent Defeat Agent: Automatic Detection of Taint-Style Vulnerabilities in LLM-based AgentsFengyu Liu, Yuan Zhang, Jiaqi Luo, Jiarun Dai 等USENIX Security 2025
- AgentInspect: Diagnosing Behavioral Failures in Artificial Intelligence AgentsRuchira Manke, Mohammad Wardat, Foutse Khomh, Hridesh RajanISSTA 2026
- CrossProbe: LLM-Empowered Cross-Project Bug Detection for Deep Learning FrameworksHao Guan, Guangdong Bai, Yepang LiuISSTA 2025 · 被引用 3 次
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test CasesZiqian Zhong, Aditi Raghunathan, Nicholas CarliniICLR 2026 · 被引用 54 次
