ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
Xin Gui, King Zhu, JinCheng Ren, Qianben Chen, Zekun Wang, Yizhi Li, Xinpeng Liu, Wenli Ren, Linyu Miao, Tianrui Qin, Ziqi Shu, He Zhu
Abstract
In recent years, the research focus of large language models (LLMs) and agents has shifted increasingly from demonstrating novel capabilities to complex reasoning and tackling challenging tasks. However, existing evaluations focus mainly on math/code contests or general tasks, while existing multi-domain academic benchmarks lack sufficient reasoning depth, leaving the field without a rigorous benchmark for high-level reasoning. To fill this gap, we introduce the ACADREASON benchmark, designed to evaluate the ability of LLMs and agents to acquire and reason over academic knowledge. It consists of 50 expert-annotated academic problems across five high-reasoning domains, including computer science, economics, law, mathematics, and philosophy. All questions are sourced from top-tier publications in recent years and undergo rigorous annotation and quality control to ensure they are both challenging and answerable. We conduct systematic evaluations over 10 mainstream LLMs and agents. The results show that most LLMs scored below 20 points, with even the cutting-edge GPT-5 achieving only 16 points. While agents achieved higher scores, none exceeded 40 points. This demonstrates the current capability gap between LLMs and agents in super-intelligent academic research tasks and highlights the challenges of ACADREASON. The code and data for the ACADREASON benchmark are available at https://github.com/OPPO-PersonalAI/Acadreason-benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 633dd57a-4d7f-47de-8979-e1fe14ddce2bBuilds on7
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian et al.NeurIPS 2025 · 354 citations
- Chain of Agents: Large Language Models Collaborating on Long-Context TasksYusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister et al.NeurIPS 2024 · 297 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
Related papers
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen et al.ICCV 2025 · 4 citations
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsXin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen et al.ICML 2025
