AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra, Nicolas Mario Baldwin, Alexis Audran-Reiss, Michael Kuchnik, Despoina Magka, Minqi Jiang, Alisia Maria Lupidi, Andrei Lupu, Roberta Raileanu
Abstract
AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6 % to 47.7 %. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning.
Finally, to conduct experiments, we develop AI Research Agent dojo (AIRA-dojo), a framework that provides a scalable and customizable environment for AI research agents. First, AIRA-dojo exposes a robust and flexible interface to compute resources, which is essential for building effective agents. The baseline, AIDE, implemented in AIRA-dojo achieves a performance increase of 10.68 % (absolute scale) over the reported results [4]. Second, AIRA-dojo enables users to experiment with custom operators, search policies, evaluation methods, and tasks within a comparable setup. This facilitates a rigorous scientific study of AI research automation. Our code is open-sourced at: https://github.com/facebookresearch/aira-dojo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b2ef90a-d7ee-45d9-9930-6513dd21b414Cited by top-tier papers9
- TwinMarket: A Scalable Behavioral and Social Simulation for Financial MarketsYuzhe Yang, Yifei Zhang, Minghao Wu, Kaidi Zhang et al.NeurIPS 2025 · 64 citations
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use InsteadFeiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica et al.ICLR 2026 · 27 citations
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen et al.ICLR 2026 · 25 citations
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&DFrancis Ward, Teun van der Weij, Hanna Gábor, Sam Martin et al.NeurIPS 2025 · 11 citations
- Can We Predict Before Executing Machine Learning Agents?Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao et al.ACL 2026 · 6 citations
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
Related papers
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin et al.NeurIPS 2025 · 58 citations
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu et al.ICLR 2026 · 35 citations
- MARS: Modular Agent with Reflective Search for Automated AI ResearchJiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng et al.ICML 2026 · 14 citations
- CoMind: Towards Community-Driven Agents for Machine Learning EngineeringSijie Li, Weiwei Sun, Shanda Li, Ameet Talwalkar et al.ICLR 2026 · 3 citations
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language AgentsQizheng Li, Yifei Zhang, Xiao Yang, Xu Yang et al.ICML 2026 · 3 citations
