MLE-Smith: Scaling MLE Tasks with Automated Multi-agent Pipeline
Rushi Qiang, Yuchen Zhuang, Anikait Singh, Percy Liang, Chao Zhang, Sherry Yang, Bo Dai
Abstract
While Language Models (LMs) have made significant progress in automating machine learning engineering (MLE), the acquisition of high-quality MLE training data is significantly constrained. Current MLE benchmarks suffer from low scalability and limited applicability because they rely on static, manually curated tasks that demand extensive time and manual effort to produce. We introduce MLE-Smith, a fully automated multi-agent pipeline, to transform raw datasets into competition-style MLE challenges through an efficient generate--verify--execute paradigm for scaling MLE tasks with verifiable quality, real-world usability and rich diversity. The proposed multi-agent pipeline in MLE-Smith drives structured task design and standardized refactoring, coupled with a hybrid verification mechanism that enforces strict structural rules and high-level semantic soundness. It further validates empirical solvability and real-world fidelity through interactive execution. We apply MLE-Smith to 224 of real-world datasets and generates 606 tasks spanning multiple categories, objectives, and modalities, demonstrating that MLE-Smith can work effectively across a wide range of real-world datasets. Evaluation on generated tasks shows that the performance of eight mainstream and cutting-edge LLMs on MLE-Smith tasks is strongly correlated with their performance on carefully human-designed tasks, highlighting the effectiveness of the MLE-Smith in scaling up MLE tasks while maintaining task quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6356cfdd-62d5-4fa1-8b11-926da4e67df9Cited by top-tier papers1
Ask how each one uses itBuilds on15
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- AdaPlanner: Adaptive Planning from Feedback with Language ModelsHaotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai et al.NeurIPS 2023 · 257 citations
Related papers
- MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software EngineeringChuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen et al.ICML 2026 · 3 citations
- Scaling Synthetic Task Generation for Agents via ExplorationRam Ramrakhya, Andrew Szot, Omar Attia, Bogdan Mazoure et al.ICLR 2026 · 15 citations
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP ToolsWenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen et al.ACL 2026 · 8 citations
- SWE-rebench V2: Language-Agnostic SWE Task Collection at ScaleIbragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Aleksandr GolubevICML 2026 · 13 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
