Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
Yujie Zhao, Lanxiang Hu, Yang Wang, Minmin Hou, Hao Zhang, Ke Ding, Jishen Zhao
摘要
Multi-Agent System (MAS) and Reinforcement Learning (RL) are both widely adopted to improve large language model (LLM) agentic performance. MAS strengthens task-specialized performance via role-based orchestration; RL leverages environment rewards to train stronger policies, such as Group Relative Policy Optimization (GRPO)-style optimization. Yet applying on-policy RL training to MAS is underexplored. While promising, it poses several challenges. On the algorithm side, Standard GRPO grouping assumptions fail in MAS because prompts differ by role and turn. On the system side, the training system needs to support MAS-workflow-based rollouts and on-policy updates for both single and multiple policy models. To address these issues, we introduce AT-GRPO, consisting of (i) an Agent- and Turn-wise grouped RL algorithm tailored for MAS and (ii) a system to support both single-policy and multi-policy training. Across game, plan, coding, and math tasks, AT-GRPO demonstrates substantial performance gains across diverse domains. Especially on long-horizon planning tasks, AT-GRPO boosts accuracy from a 14.0–47.0% single-agent RL baseline to 96.0–99.5%. Furthermore, it improves reasoning performance, with an average gain of 3.87–7.62% on coding and 9.0-17.93% on math. The code are available at https://github.com/pettingllms-ai/PettingLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Decentralized LLM Collaboration with Multi-Agent Actor CriticShuo Liu, Tianle Chen, Ryan Amiri, Christopher AmatoICML 2026 · 被引用 6 次
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 被引用 3 次
它引用的顶会 Paper14
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionWei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang 等NeurIPS 2024 · 被引用 210 次
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningBo Liu, Simon Yu, Zichen Liu, Leon Guertler 等ICLR 2026 · 被引用 88 次
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement LearningChanwoo Park, Seungju Han, Xingzhi Guo, Asuman E. Ozdaglar 等ACL 2025 · 被引用 63 次
相关 Paper
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan 等ACL 2026 · 被引用 5 次
- LLM Collaboration with Multi-Agent Reinforcement LearningShuo Liu, Zeyu Liang, Xueguang Lyu, Christopher AmatoAAAI 2026
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen 等ICML 2026 · 被引用 24 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han 等ICLR 2026 · 被引用 47 次
