Characterizing Deep Research: A Benchmark and Formal Definition
Abhinav Java, Ashmit Khandelwal, Sukruta Prakash Midigeshi, Aaron Halfaker, Amit Jayant Deshpande, Navin Goyal, Ankur Gupta, Nagarajan Natarajan, Amit Sharma
摘要
Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of deep research -- a term also adopted by recent models targeting these capabilities. Despite growing interest, the scope of the deep research task remains underdefined and its distinction from other reasoning-intensive problems is poorly understood. In this paper, we propose a formal characterization of the deep research (DR) task and introduce a benchmark to evaluate the performance of DR systems. We argue that the core defining feature of deep research is not the production of lengthy report-style outputs, but rather the high fan-out over concepts required during the search process, i.e., broad and reasoning-intensive exploration. To enable objective evaluation, we define DR using an intermediate output representation that encodes key claims uncovered during search-separating the reasoning challenge from surface-level report generation. Based on this formulation, we propose a diverse, challenging benchmark LiveDRBench with 100 challenging tasks over scientific topics (e.g., datasets, materials discovery, prior art search) and public interest events (e.g., flight incidents, movie awards). Across state-of-the-art DR systems, F1 score ranges between 0.02 and 0.72 for any sub-category. OpenAI's model performs the best with an overall F1 score of 0.55. Analysis of reasoning traces reveals the distribution over the number of referenced sources, branching, and backtracking events executed by current DR systems, motivating future directions for improving their search mechanisms and grounding capabilities. The benchmark is available at https://github.com/microsoft/LiveDRBench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang 等ICLR 2026 · 被引用 83 次
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox 等ICLR 2026 · 被引用 18 次
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report GenerationJanghoon Han, Heegyu Kim, Changho Lee, Dahm Lee 等ICML 2026 · 被引用 9 次
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman 等ACL 2026 · 被引用 1 次
- AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang 等ACL 2026
它引用的顶会 Paper7
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- Precise Information Control in Long-Form Text GenerationJacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li 等NeurIPS 2025 · 被引用 8 次
- VeriTrail: Closed-Domain Hallucination Detection with TraceabilityDasha Metropolitansky, Jonathan LarsonICLR 2026 · 被引用 3 次
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi 等ICLR 2025
相关 Paper
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen 等ICLR 2026 · 被引用 37 次
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas 等ICLR 2026 · 被引用 1 次
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsWei Liu, Peijie Yu, Michele Orini, Yali Du 等ICML 2026 · 被引用 2 次
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 被引用 2 次
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian 等NeurIPS 2025 · 被引用 354 次
