Investigating Component Contributions in Multi-Agent ML Systems
Junsung Kim, Ilia Mireskandari, Seungwan Son, Yifan Zhou, Khizer Shahid, Dylan Dai
摘要
Autonomous agents for machine learning engineering have advanced rapidly, yet comparing their effectiveness remains difficult. Existing systems combine different techniques-multiagent decomposition, iterative refinement, memory management, and planning-in varying configurations, making it unclear which components actually drive performance. Complicating evaluation, existing benchmarks rely on historical competitions whose data likely contaminates LLM training corpora and whose static baselines reflect outdated human performance. To address this, we conduct approximately 4,000 controlled experiments systematically ablating architectural components, alongside K-LIVE 1 a new benchmark of 25 competitions that provides a dynamic evaluation environment with minimal data contamination. Our findings challenge common design assumptions: in our evaluation, iterative feedback contributes more than architectural complexity, and fixed-role multi-agent coordination consistently underperforms a single-agent baseline. These results provide concrete guidance for practitioners building ML engineering agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 被引用 209 次
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningSiyuan Guo, Cheng Deng, Ying Wen, Hechang Chen 等ICML 2024 · 被引用 107 次
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung 等ICLR 2025 · 被引用 9 次
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu 等ICLR 2025 · 被引用 7 次
相关 Paper
- CoMind: Towards Community-Driven Agents for Machine Learning EngineeringSijie Li, Weiwei Sun, Shanda Li, Ameet Talwalkar 等ICLR 2026 · 被引用 3 次
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin 等NeurIPS 2025 · 被引用 58 次
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language AgentsQizheng Li, Yifei Zhang, Xiao Yang, Xu Yang 等ICML 2026 · 被引用 3 次
- MARS: Modular Agent with Reflective Search for Automated AI ResearchJiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng 等ICML 2026 · 被引用 14 次
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang 等ACL 2025 · 被引用 97 次
