Demystifying Deep Search: A Holistic Evaluation with Hint-free Multi-Hop Questions and Factorised Metrics
Maojia Song, Renhang Liu, Xinyu Wang, Yong Jiang, Pengjun Xie, Fei Huang, Soujanya Poria, Jingren Zhou
摘要
RAG (Retrieval-Augmented Generation) systems and web agents are increasingly evaluated on multi-hop deep search tasks, yet current practice suffers from two major limitations. First, most benchmarks leak the reasoning path in the question text, allowing models to follow surface cues rather than discover reasoning chains autonomously. Second, evaluation is typically reduced to a single pass rate, which collapses diverse behaviours into one score and obscures whether failures stem from inadequate search, poor knowledge use, or inappropriate refusal. To address these issues, we present WebDetective, a benchmark of hint-free multi-hop questions paired with a controlled Wikipedia sandbox that ensures full traceability of model actions, and a holistic evaluation framework that separates search sufficiency, knowledge utilisation, and refusal behaviour. Our evaluation of 25 state-of-the-art models reveals systematic weaknesses across all architectures: models struggle with knowledge utilisation despite having sufficient evidence and demonstrate near-absent appropriate refusal when evidence is lacking. These patterns expose a fundamental gap: today's systems excel at executing given reasoning paths but fail when required to discover them. We develop an agentic workflow, EvidenceLoop, that explicitly targets the challenges our benchmark identifies, incorporating verification loops and systematic evidence tracking that improve both search and synthesis capabilities. This baseline demonstrates that WebDetective's diagnostic framework can guide concrete architectural improvements, establishing our benchmark as a critical tool for developing genuinely autonomous reasoning systems rather than pattern-following agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationZhengwei Tao, Jialong Wu, Wenbiao Yin, Pu Wu 等ICLR 2026 · 被引用 115 次
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen 等ICLR 2026 · 被引用 66 次
- Search-o1: Agentic Search-Enhanced Large Reasoning ModelsXiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang 等EMNLP 2025 · 被引用 12 次
相关 Paper
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning ChainsXuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai 等ICLR 2026 · 被引用 14 次
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QAMinghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang 等ACL 2026 · 被引用 1 次
- WebWalker: Benchmarking LLMs in Web TraversalJialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang 等ACL 2025
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalJiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang 等CVPR 2026 · 被引用 2 次
- TRACE: Traversal Retrieval-Augmented Chain of Evidence for Document UnderstandingLiqi He, Zuchao Li, Hao Huang, Ping WangACL 2026
