Investigating Component Contributions in Multi-Agent ML Systems
Junsung Kim, Ilia Mireskandari, Seungwan Son, Yifan Zhou, Khizer Shahid, Dylan Dai
Abstract
Autonomous agents for machine learning engineering have advanced rapidly, yet comparing their effectiveness remains difficult. Existing systems combine different techniques-multiagent decomposition, iterative refinement, memory management, and planning-in varying configurations, making it unclear which components actually drive performance. Complicating evaluation, existing benchmarks rely on historical competitions whose data likely contaminates LLM training corpora and whose static baselines reflect outdated human performance. To address this, we conduct approximately 4,000 controlled experiments systematically ablating architectural components, alongside K-LIVE 1 a new benchmark of 25 competitions that provides a dynamic evaluation environment with minimal data contamination. Our findings challenge common design assumptions: in our evaluation, iterative feedback contributes more than architectural complexity, and fixed-role multi-agent coordination consistently underperforms a single-agent baseline. These results provide concrete guidance for practitioners building ML engineering agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c283442-c4e6-4bc3-aa01-c1390c679b14Builds on6
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningSiyuan Guo, Cheng Deng, Ying Wen, Hechang Chen et al.ICML 2024 · 107 citations
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung et al.ICLR 2025 · 9 citations
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu et al.ICLR 2025 · 7 citations
Related papers
- CoMind: Towards Community-Driven Agents for Machine Learning EngineeringSijie Li, Weiwei Sun, Shanda Li, Ameet Talwalkar et al.ICLR 2026 · 3 citations
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin et al.NeurIPS 2025 · 58 citations
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language AgentsQizheng Li, Yifei Zhang, Xiao Yang, Xu Yang et al.ICML 2026 · 3 citations
- MARS: Modular Agent with Reflective Search for Automated AI ResearchJiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng et al.ICML 2026 · 14 citations
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang et al.ACL 2025 · 97 citations
