ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
Zexi Liu, Jingyi Chai, Xinyu Zhu, shuo tang, Rui Ye, Weiyu Ma, Bo Zhang, LEI BAI, Siheng Chen
摘要
The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering. However, the dominant prompt-based paradigm exhibits limitations: smaller models lack the capacity to learn from execution trajectories for generalization, while large proprietary models incur high computational overhead, restricting accessibility and scalability. Focusing on this, for the first time, we explore the paradigm of learning-based agentic ML, where an LLM agent learns through interactive experimentation on ML tasks using online reinforcement learning (RL). To realize this, we propose a novel agentic ML training framework with three key components: (1) exploration-enriched fine-tuning, which enables LLM agents to generate diverse actions for enhanced RL exploration; (2) step-wise RL, which enables training on a single action step, accelerating experience collection and improving training efficiency; (3) an agentic ML-specific reward module, which unifies varied ML feedback signals into consistent rewards for RL optimization. Leveraging this framework, we train ML-Agent, driven by a 7B-sized Qwen-2.5 LLM for autonomous ML. Despite training on only 9 ML tasks, our 7B-sized ML-Agent achieves comparable performance to agents using much larger proprietary LLMs (e.g., GPT-5) but at significantly lower computational cost, demonstrating strong performance and cross-task generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement LearningWei Yang, Jesse ThomasonAAAI 2026 · 被引用 9 次
- Can We Predict Before Executing Machine Learning Agents?Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao 等ACL 2026 · 被引用 6 次
- Hybrid-Gym: Training Coding Agents to Generalize Across TasksYiqing Xie, Emmy Liu, Gaokai Zhang, Nachiket Kotalwar 等ICML 2026 · 被引用 4 次
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context ManagementZherui Yang, Fan Liu, Yansong Ning, Hao LiuKDD 2026 · 被引用 3 次
- MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference LearningNian Ran, Zhongzheng Li, Yue Wang, Qingsong Ran 等ICML 2026
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 被引用 209 次
- An ADMM Based Framework for AutoML Pipeline ConfigurationSijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf 等AAAI 2020 · 被引用 82 次
相关 Paper
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
- Offline Training of Language Model Agents with Functions as Learnable WeightsShaokun Zhang, Jieyu Zhang, Jiale Liu, Linxin Song 等ICML 2024 · 被引用 41 次
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang 等ICLR 2026 · 被引用 26 次
- MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and InferenceKaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng 等ICLR 2026
