AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun
Abstract
Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods often renders reproduction a laborintensive endeavor, necessitating profound domain expertise. To address this, we introduce the paper lineage, which systematically mines implicit knowledge from the cited literature. This algorithm serves as the backbone of our proposed AUTOREPRODUCE, a multi-agent framework designed to autonomously reproduce experimental code in a complete, endto-end manner. To ensure code executability, AUTOREPRODUCE incorporates a samplingbased unit testing strategy for rapid validation. To assess reproduction capabilities, we introduce REPRODUCEBENCH, a benchmark featuring verified implementations, alongside comprehensive metrics for evaluating both reproduction and execution fidelity. Extensive evaluations on PaperBench and REPRO-DUCEBENCH demonstrate that AUTOREPRO-DUCE consistently surpasses existing baselines across all metrics. Notably, it yields substantial improvements in reproduction fidelity and final execution performance. The code is available at https://github.com/AI9Stars/ AutoReproduce .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c981a53-1146-4db7-82b4-a7fdc79b5e76Cited by top-tier papers4
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket et al.ICLR 2026 · 51 citations
- SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?Udari Sehwag, Elaine Lau, Haniyeh Oskouie, Shayan Shabihi et al.ICML 2026 · 1 citation
- NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into CodeSeemandhar Jain, Keshav Gupta, Kunal Gupta, Manmohan ChandrakerCVPR 2026 · 1 citation
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu et al.SIGCOMM 2026
Builds on17
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- IEBins: Iterative Elastic Bins for Monocular Depth EstimationShuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu et al.NeurIPS 2023 · 114 citations
Related papers
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code MaskingGyeongwon James Kim, Alex Wilf, Louis-Philippe Morency, Daniel FriedICLR 2026 · 12 citations
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang et al.EMNLP 2025
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
- PaperBench: Evaluating AI's Ability to Replicate AI ResearchGiulio Starace, Oliver Jaffe, Dane Sherburn, James Aung et al.ICML 2025
- AI-Researcher: Autonomous Scientific InnovationJiabin Tang, Lianghao Xia, Zhonghang Li, Chao HuangNeurIPS 2025 · 101 citations
