DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Mingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang, Xiaorui Wang, Zhendong Mao
Abstract
Deep Research Agents (DRAs) are emerging as one of the most practical classes of LLM-based agents. Given an open-ended research task, they find, analyze, and synthesize large numbers of online sources to produce a comprehensive report at the level of a research analyst. This can compress hours of manual desk research into minutes. However, a comprehensive benchmark for systematically evaluating the capabilities of these agents remains absent. To bridge this gap, we introduce DeepResearch Bench, a benchmark consisting of 100 PhD-level research tasks, each meticulously crafted by domain experts across 22 distinct fields. To evaluate DRAs comprehensively, we propose two complementary and fully automated methodologies. The first is a reference-based method with adaptive criteria to assess the quality of generated research reports. The second evaluates a DRA’s information‑retrieval and collection capabilities by assessing its effective citation count and overall citation accuracy. By conducting extensive human consistency experiments, we demonstrate that our evaluation methods are highly aligned with expert judges and faithfully reflect human judgments of quality differences among DRA-generated content. We are open-sourcing DeepResearch Bench and key components of these frameworks to accelerate the development of practical LLM-based agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b843551-b63b-4113-a3ac-9cf3f8e3d76eCited by top-tier papers58
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju et al.ICLR 2026 · 109 citations
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison et al.ICML 2026 · 78 citations
- MemEvolve: Meta-Evolution of Agent Memory SystemsGuibin Zhang, Haotian Ren, Chong Zhan, Junhao Wang et al.ICML 2026 · 69 citations
Builds on12
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian et al.NeurIPS 2025 · 354 citations
Related papers
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang et al.ICLR 2026 · 13 citations
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox et al.ICLR 2026 · 18 citations
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen et al.ICLR 2026 · 37 citations
- Towards Knowledgeable Deep Research: Framework and BenchmarkWenxuan Liu, Zixuan Li, Long Bai, Chunmao Zhang et al.SIGIR 2026
