Multi-Agent Teams Hold Experts Back
Aneesh Pappu, Batu El, Hancheng Cao, Carmelo di Nolfo, Yanchao Sun, Meng Cao, James Zou
Abstract
Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, effective coordination cannot be fully designed in advance and must instead emerge through interaction. However, most prior work enforces coordination through fixed roles, workflows, or aggregation rules, leaving open the question of how well self-organizing teams perform when coordination is unconstrained. Drawing on organizational psychology, we study whether self-organizing LLM teams achieve strong synergy , where team performance matches or exceeds the best individual member. Across human-inspired and frontier ML benchmarks, we find that---unlike human teams---LLM teams consistently fail to match their expert agent's performance, even when explicitly told who the expert is, incurring performance losses of up to 41.1% on ML benchmarks. Decomposing this failure, we show that expert leveraging, rather than identification, is the primary bottleneck. Conversational analysis reveals a tendency toward integrative compromise---averaging expert and non-expert views rather than appropriately weighting expertise---which increases with team size and correlates negatively with performance. Interestingly, this consensus-seeking behavior improves robustness to adversarial agents, suggesting a trade-off between alignment and effective expertise utilization. Our findings reveal a significant gap in the ability of self-organizing multi-agent teams to harness the collective expertise of their members.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09af634f-719c-4793-a36c-2babf08abc7eBuilds on7
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent SystemsYingxuan Yang, Huacan Chai, Shuai Shao, Yuanyi Song et al.NeurIPS 2025 · 104 citations
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?Hyeong Kyu Choi, Xiaojin Zhu, Sharon LiNeurIPS 2025 · 93 citations
- Latent Collaboration in Multi-Agent SystemsJiaru Zou, Xiyuan Yang, Ruizhong Qiu, Gaotang Li et al.ICML 2026 · 42 citations
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMsYuxuan Li, Aoi Naito, Hirokazu ShiradoICML 2026 · 9 citations
Related papers
- More Capable, Less Cooperative? When LLMs Fail at Zero-Cost CollaborationAdvait Yadav, Sidney Black, Oliver SourbutICML 2026 · 2 citations
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang et al.ICLR 2025
- SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM SystemsYuzhe Zhang, Feiran Liu, Yi Shan, Xinyi Huang et al.ACL 2026 · 5 citations
- Advancing Collaborative Debates with Role Differentiation through Multi-Agent Reinforcement LearningHaoran Li, Ziyi Su, Yun Xue, Zhiliang Tian et al.ACL 2025 · 8 citations
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun et al.ICLR 2026 · 22 citations
