Deep Research Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
Haiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu, Jiaxuan Lu, Di Yu, Jianbao Cao, Ben Gao, Jiaqing Xie, Aoran Wang, Wenlong Zhang, Philip Torr, Dongzhan Zhou
Abstract
Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating their research capability faithfully is rather challenging due to the difficulty of collecting frontier research questions that genuinely capture researchers’ attention and intellectual curiosity. To address this gap, we introduce DeepResearch Arena, a benchmark grounded in academic seminars that capture rich expert discourse and interaction, better reflecting real-world research environments and reducing the risk of data leakage. To automatically construct DeepResearch Arena, we propose a Multi-Agent Hierarchical Task Generation (MAHTG) system that extracts research-worthy inspirations from seminar transcripts. The MAHTG system further translates research-worthy inspirations into high-quality research tasks, ensuring the traceability of research task formulation while filtering noise. With the MAHTG system, we curate DeepResearch Arena with over 10,000 high-quality research tasks from over 200 academic seminars, spanning 12 disciplines, such as literature, history, and science. Our extensive evaluation shows that DeepResearch Arena presents substantial challenges for current state-of-the-art agents, with clear performance gaps observed across different models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e78c1a3b-c7cc-4709-a331-b4279f62cdf1Cited by top-tier papers10
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report GenerationJanghoon Han, Heegyu Kim, Changho Lee, Dahm Lee et al.ICML 2026 · 9 citations
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?Junchi Yu, Yujie Liu, Jindong Gu, Philip H. S. Torr et al.NeurIPS 2025 · 8 citations
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 2 citations
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsWei Liu, Peijie Yu, Michele Orini, Yali Du et al.ICML 2026 · 2 citations
Builds on7
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu et al.ACL 2025 · 88 citations
- Thought Propagation: an Analogical Approach to Complex Reasoning with Large Language ModelsJunchi Yu, Ran He, Zhitao YingICLR 2024 · 44 citations
Related papers
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- Characterizing Deep Research: A Benchmark and Formal DefinitionAbhinav Java, Ashmit Khandelwal, Sukruta Prakash Midigeshi, Aaron Halfaker et al.ICLR 2026 · 30 citations
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen et al.ICLR 2026 · 37 citations
- Towards Knowledgeable Deep Research: Framework and BenchmarkWenxuan Liu, Zixuan Li, Long Bai, Chunmao Zhang et al.SIGIR 2026
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox et al.ICLR 2026 · 18 citations
