InteractComp: Evaluating Search Agents With Ambiguous Queries
Mingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong, Jiayi Zhang, Fashen Ren, Jinyi Bai, Fuzhen Yang, Dayi Miao, Zhaoyang Yu, Yifan WU, Yanfei Zhang
Abstract
Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failure mode: agents may face ambiguous requests where the intended target cannot be identified without clarification. Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce InteractComp, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. Following the principle of easy to verify, interact to disambiguate, we construct 210 expert-curated questions across 9 domains through a target-distractor methodology that creates controlled ambiguity resolvable only through interaction. Evaluation of 17 models reveals striking failure: the best model achieves only 13.73% accuracy despite 71.50% with complete context, exposing systematic overconfidence rather than reasoning deficits. Forced interaction produces dramatic gains, demonstrating latent capability current strategies fail to engage. Longitudinal analysis shows interaction capabilities stagnated over 15 months while search performance improved seven-fold, revealing a critical blind spot. This stagnation, coupled with the immediate feedback inherent to search tasks, makes InteractComp a valuable resource for both evaluating and training interaction capabilities in search agents. The code is available at https://github.com/FoundationAgents/InteractComp
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Atom of Thoughts for Markov LLM Test-Time ScalingFengwei Teng, Quan Shi, Zhaoyang Yu, Jiayi Zhang et al.NeurIPS 2025 · 73 citations
- DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL FrameworkBoyan Li, Chong Chen, Zhujun Xue, Yinan Mei et al.SIGMOD 2026 · 40 citations
Related papers
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu et al.ICLR 2026 · 29 citations
- BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search AgentsZijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie et al.ACL 2026
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren et al.ACL 2025
- InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed InformationJiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang et al.ICML 2026
- Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software EngineeringSanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap et al.ICLR 2026 · 35 citations
