WideSearch: Benchmarking Agentic Broad Info-Seeking
Ryan Wong, Jiawei Wang, Junjie Zhao, Li Chen, Yan Gao, Long Zhang, Xuan Zhou, Zuo Wang, Kai Xiang, Ge Zhang, Wenhao Huang, Yang Wang, Ke Wang
摘要
From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberate humans from this tedious work. However, the capability of these agents to perform such"wide-context"collection reliably and completely remains largely unevaluated due to a lack of suitable benchmarks. To bridge this gap, we introduce WideSearch, a new benchmark engineered to evaluate agent reliability on these large-scale collection tasks. The benchmark features 200 manually curated questions (100 in English, 100 in Chinese) from over 15 diverse domains, grounded in real user queries. Each task requires agents to collect large-scale atomic information, which could be verified one by one objectively, and arrange it into a well-organized output. A rigorous five-stage quality control pipeline ensures the difficulty, completeness, and verifiability of the dataset. We benchmark over 10 state-of-the-art agentic search systems, including single-agent, multi-agent frameworks, and end-to-end commercial systems. Most systems achieve overall success rates near 0%, with the best performer reaching just 5%. However, given sufficient time, cross-validation by multiple human testers can achieve a near 100% success rate. These results demonstrate that present search agents have critical deficiencies in large-scale information seeking, underscoring urgent areas for future research and development in agentic search. Our dataset, evaluation pipeline, and benchmark results have been publicly released at https://widesearch-seed.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Scaling Long-Horizon Agent via Context FoldingWeiwei Sun, Lu Miao, Zhan Ling, Kang Liu 等ICML 2026 · 被引用 104 次
- Demystifying Deep Search: A Holistic Evaluation with Hint-free Multi-Hop Questions and Factorised MetricsMaojia Song, Renhang Liu, Xinyu Wang, Yong Jiang 等ICLR 2026 · 被引用 7 次
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang 等ACL 2026 · 被引用 4 次
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 被引用 2 次
- UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information SeekingChang Liu, Chuqiao Kuang, Tianyi Zhuang, Yuxin Cheng 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper3
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao 等ICLR 2025
相关 Paper
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 被引用 2 次
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas 等ICLR 2026 · 被引用 1 次
- LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context GrowthWeihao Zeng, Yuzhen Huang, Junxian HeICML 2026 · 被引用 11 次
- AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic WebShanshan Zhong, Kate Shen, Chenyan XiongICML 2026 · 被引用 2 次
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou 等ACL 2026 · 被引用 1,216 次
