SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
Ben Bogin, Kejuan Yang, Shashank Gupta, Kyle Richardson, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot
摘要
Given that Large Language Models (LLMs) have made significant progress in writing code, can they now be used to autonomously reproduce results from research repositories? Such a capability would be a boon to the research community, helping researchers validate, understand, and extend prior work. To advance towards this goal, we introduce SUPER, the first benchmark designed to evaluate the capability of LLMs in setting up and executing tasks from research repositories. SUPER aims to capture the realistic challenges faced by researchers working with Machine Learning (ML) and Natural Language Processing (NLP) research repositories. Our benchmark comprises three distinct problem sets: 45 endto-end problems with annotated expert solutions, 152 sub-problems derived from the expert set that focus on specific challenges (e.g., configuring a trainer), and 604 automatically generated problems for larger-scale development. We introduce various evaluation measures to assess both task success and progress, utilizing gold solutions when available or approximations otherwise. We show that stateof-the-art approaches struggle to solve these problems with the best model (GPT-4o) solving only 16.3% of the end-to-end set, and 46.1% of the scenarios. This illustrates the challenge of this task, and suggests that SUPER can serve as a valuable resource for the community to make and measure progress. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket 等ICLR 2026 · 被引用 51 次
- LLM Agents Making Agent ToolsGeorg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic 等ACL 2025 · 被引用 41 次
- AutoReproduce: Automatic AI Experiment Reproduction with Paper LineageXuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi 等ACL 2026 · 被引用 18 次
- NetArena: Dynamic Benchmarks for AI Agents in Network AutomationYajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi 等ICLR 2026 · 被引用 17 次
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code MaskingGyeongwon James Kim, Alex Wilf, Louis-Philippe Morency, Daniel FriedICLR 2026 · 被引用 12 次
它引用的顶会 Paper21
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
相关 Paper
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang 等EMNLP 2025
- RExBench: Can coding agents autonomously implement AI research extensions?Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin 等ACL 2026 · 被引用 10 次
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun 等AAAI 2026 · 被引用 10 次
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan 等ICML 2026 · 被引用 31 次
