Multi-Agent Collaboration via Evolving Orchestration
Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang, Xiaoyin Che, Ye Tian, Xuantang Xiong, Lei Han
Abstract
Large language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organizational structures that struggle to adapt as task complexity and agent numbers grow, resulting in coordination overhead and inefficiencies. To this end, we propose a puppeteer-style paradigm for LLM-based multi-agent collaboration, where a centralized orchestrator ("puppeteer") dynamically directs agents ("puppets") in response to evolving task states. This orchestrator is trained via reinforcement learning to adaptively sequence and prioritize agents, enabling flexible and evolvable collective reasoning. Experiments on closed- and open-domain scenarios show that this method achieves superior performance with reduced computational costs. Analyses further reveal that the key improvements consistently stem from the emergence of more compact, cyclic reasoning structures under the orchestrator's evolution. Our code is available at https://github.com/OpenBMB/ChatDev/tree/puppeteer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb33e468-6f18-4bcf-bda9-ee6bfb03e3b1Cited by top-tier papers10
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun et al.ICLR 2026 · 22 citations
- MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled BenchmarksZixuan Ke, Yifei Ming, Austin Xu, Ryan Chin et al.ICML 2026 · 15 citations
- Don't Throw Away Your Pretrained ModelShangbin Feng, Wenhao Yu, Yike Wang, Hongming Zhang et al.ICLR 2026 · 10 citations
- InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual SearchKaican Li, Lewei Yao, Jiannan Wu, Tiezheng YU et al.ICLR 2026 · 10 citations
- Eigen-Agent: Adaptive Multi-Agent Scientific Reasoning with Monitor-Based RAGXiangru Tang, Wanghan Xu, Yujie Wang, Zijie Guo et al.ICLR 2026 · 8 citations
Builds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu et al.NeurIPS 2024 · 54 citations
- Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent GamesYidong He, Yutao Lai, Pengxu Yang, Jiarui Gan et al.ICML 2026
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool OrchestrationHongjin SU, Shizhe Diao, Ximing Lu, Mingjie Liu et al.ICML 2026 · 34 citations
- MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt OptimizationJian Zhang, Zhangqi Wang, Haiping Zhu, Kangda Cheng et al.AAAI 2026 · 9 citations
- BOAD: Discovering Hierarchical Software Engineering Agents via Bandit OptimizationIris Xu, Guangtao Zeng, Zexue He, Charles Jin et al.ICLR 2026 · 5 citations
