AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun
摘要
Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods often renders reproduction a laborintensive endeavor, necessitating profound domain expertise. To address this, we introduce the paper lineage, which systematically mines implicit knowledge from the cited literature. This algorithm serves as the backbone of our proposed AUTOREPRODUCE, a multi-agent framework designed to autonomously reproduce experimental code in a complete, endto-end manner. To ensure code executability, AUTOREPRODUCE incorporates a samplingbased unit testing strategy for rapid validation. To assess reproduction capabilities, we introduce REPRODUCEBENCH, a benchmark featuring verified implementations, alongside comprehensive metrics for evaluating both reproduction and execution fidelity. Extensive evaluations on PaperBench and REPRO-DUCEBENCH demonstrate that AUTOREPRO-DUCE consistently surpasses existing baselines across all metrics. Notably, it yields substantial improvements in reproduction fidelity and final execution performance. The code is available at https://github.com/AI9Stars/ AutoReproduce .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket 等ICLR 2026 · 被引用 51 次
- SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?Udari Sehwag, Elaine Lau, Haniyeh Oskouie, Shayan Shabihi 等ICML 2026 · 被引用 1 次
- NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into CodeSeemandhar Jain, Keshav Gupta, Kunal Gupta, Manmohan ChandrakerCVPR 2026 · 被引用 1 次
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu 等SIGCOMM 2026
它引用的顶会 Paper17
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- IEBins: Iterative Elastic Bins for Monocular Depth EstimationShuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 等NeurIPS 2023 · 被引用 114 次
相关 Paper
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code MaskingGyeongwon James Kim, Alex Wilf, Louis-Philippe Morency, Daniel FriedICLR 2026 · 被引用 12 次
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang 等EMNLP 2025
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 被引用 86 次
- PaperBench: Evaluating AI's Ability to Replicate AI ResearchGiulio Starace, Oliver Jaffe, Dane Sherburn, James Aung 等ICML 2025
- AI-Researcher: Autonomous Scientific InnovationJiabin Tang, Lianghao Xia, Zhonghang Li, Chao HuangNeurIPS 2025 · 被引用 101 次
