RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts
Hjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Joshua Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia
摘要
Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic and have a direct comparison to human performance. We introduce v1), which consists of 7 challenging, openended ML research engineering environments and data from 71 8-hour attempts by 61 distinct human experts. We confirm that our experts make progress in the environments given 8 hours, with 82% of expert attempts achieving a non-zero score and 24% matching or exceeding our strong reference solutions. We compare humans to several public frontier models through best-of-k with varying time budgets and agent designs, and find that the best AI agents achieve a score 4× higher than human experts when both are given a total time budget of 2 hours per environment. However, humans currently display better returns to increasing time budgets, narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2× the score of the top AI agent when both are given 32 total hours (across different attempts). Qualitatively, we find that modern AI agents possess significant expertise in many ML topics-e.g. an agent wrote a faster custom Triton kernel than any of our human experts'-and can generate and test solutions over ten times faster than humans, at much lower cost. We open-source the evaluation environments, human expert data, analysis code and agent trajectories to facilitate future research. 1 * Qally's. Work done in collaboration with METR. † Ordered alphabetically. ‡ Redwood Research. Work done while at METR. § Independent. Work done in collaboration with METR. ¶ Harvard University. Work done while at METR. 1 Environments can be found at github.com/METR/ai-rd-tasks and agent trajectories can be found at transcripts.metr.org. Analysis code and anonymized human expert data coming soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng 等NeurIPS 2025 · 被引用 160 次
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra 等NeurIPS 2025 · 被引用 71 次
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- NetArena: Dynamic Benchmarks for AI Agents in Network AutomationYajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi 等ICLR 2026 · 被引用 17 次
- Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic TasksTajamul Ashraf, Amal Saqib, Hanan Gani, Muhra AlMahri 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
相关 Paper
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim 等ICLR 2026 · 被引用 154 次
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao 等ICML 2026
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?Ben Rank, Hardik Bhatnagar, Ameya Pandurang Prabhu, Shira Eisenberg 等ICML 2026 · 被引用 28 次
- BRIDGE: Predicting Human Task Completion Time From Model PerformanceFengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau 等ICML 2026 · 被引用 5 次
- InnovatorBench: Evaluating Agents' Ability to Conduct Innovative AI ResearchYunze Wu, Dayuan Fu, Weiye Si, Zhen Huang 等ICLR 2026 · 被引用 9 次
