Lune

ICML2026Top-tier venue

Investigating Component Contributions in Multi-Agent ML Systems

Junsung Kim, Ilia Mireskandari, Seungwan Son, Yifan Zhou, Khizer Shahid, Dylan Dai

2026Year

Abstract

Autonomous agents for machine learning engineering have advanced rapidly, yet comparing their effectiveness remains difficult. Existing systems combine different techniques-multiagent decomposition, iterative refinement, memory management, and planning-in varying configurations, making it unclear which components actually drive performance. Complicating evaluation, existing benchmarks rely on historical competitions whose data likely contaminates LLM training corpora and whose static baselines reflect outdated human performance. To address this, we conduct approximately 4,000 controlled experiments systematically ablating architectural components, alongside K-LIVE 1 a new benchmark of 25 competitions that provides a dynamic evaluation environment with minimal data contamination. Our findings challenge common design assumptions: in our evaluation, iterative feedback contributes more than architectural complexity, and fixed-role multi-agent coordination consistently underperforms a single-agent baseline. These results provide concrete guidance for practitioners building ML engineering agents.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6c283442-c4e6-4bc3-aa01-c1390c679b14

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines