WideSearch: Benchmarking Agentic Broad Info-Seeking
Ryan Wong, Jiawei Wang, Junjie Zhao, Li Chen, Yan Gao, Long Zhang, Xuan Zhou, Zuo Wang, Kai Xiang, Ge Zhang, Wenhao Huang, Yang Wang, Ke Wang
Abstract
From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberate humans from this tedious work. However, the capability of these agents to perform such"wide-context"collection reliably and completely remains largely unevaluated due to a lack of suitable benchmarks. To bridge this gap, we introduce WideSearch, a new benchmark engineered to evaluate agent reliability on these large-scale collection tasks. The benchmark features 200 manually curated questions (100 in English, 100 in Chinese) from over 15 diverse domains, grounded in real user queries. Each task requires agents to collect large-scale atomic information, which could be verified one by one objectively, and arrange it into a well-organized output. A rigorous five-stage quality control pipeline ensures the difficulty, completeness, and verifiability of the dataset. We benchmark over 10 state-of-the-art agentic search systems, including single-agent, multi-agent frameworks, and end-to-end commercial systems. Most systems achieve overall success rates near 0%, with the best performer reaching just 5%. However, given sufficient time, cross-validation by multiple human testers can achieve a near 100% success rate. These results demonstrate that present search agents have critical deficiencies in large-scale information seeking, underscoring urgent areas for future research and development in agentic search. Our dataset, evaluation pipeline, and benchmark results have been publicly released at https://widesearch-seed.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa76eced-09a6-460c-ba25-997bed5e213dCited by top-tier papers10
- Scaling Long-Horizon Agent via Context FoldingWeiwei Sun, Lu Miao, Zhan Ling, Kang Liu et al.ICML 2026 · 104 citations
- Demystifying Deep Search: A Holistic Evaluation with Hint-free Multi-Hop Questions and Factorised MetricsMaojia Song, Renhang Liu, Xinyu Wang, Yong Jiang et al.ICLR 2026 · 7 citations
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang et al.ACL 2026 · 4 citations
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 2 citations
- UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information SeekingChang Liu, Chuqiao Kuang, Tianyi Zhuang, Yuxin Cheng et al.ICLR 2026 · 1 citation
Builds on3
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao et al.ICLR 2025
Related papers
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context GrowthWeihao Zeng, Yuzhen Huang, Junxian HeICML 2026 · 11 citations
- AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic WebShanshan Zhong, Kate Shen, Chenyan XiongICML 2026 · 2 citations
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
