Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, Yaodong Yang
Abstract
Trust region methods rigorously enabled reinforcement learning (RL) agents to learn monotonically improving policies, leading to superior performance on a variety of tasks. Unfortunately, when it comes to multi-agent reinforcement learning (MARL), the property of monotonic improvement may not simply apply; this is because agents, even in cooperative games, could have conflicting directions of policy updates. As a result, achieving a guaranteed improvement on the joint policy where each agent acts individually remains an open challenge. In this paper, we extend the theory of trust region learning to cooperative MARL. Central to our findings are the multi-agent advantage decomposition lemma and the sequential policy update scheme. Based on these, we develop Heterogeneous-Agent Trust Region Policy Optimisation (HATPRO) and Heterogeneous-Agent Proximal Policy Optimisation (HAPPO) algorithms. Unlike many existing MARL algorithms, HATRPO/HAPPO do not need agents to share parameters, nor do they need any restrictive assumptions on decomposibility of the joint value function. Most importantly, we justify in theory the monotonic improvement property of HATRPO/HAPPO. We evaluate the proposed methods on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that HATRPO and HAPPO significantly outperform strong baselines such as IPPO, MAPPO and MADDPG on all tested tasks, thereby establishing a new state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19d51880-1c6f-4e41-8012-1b4bee073920Cited by top-tier papers73
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Heterogeneous Agent Q-weighted Policy OptimizationBor-Jiun Lin, Chun-Yi LeeICLR 2026 · 102 citations
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang et al.NeurIPS 2022 · 65 citations
- Learning Multi-Agent Communication from Graph Modeling PerspectiveShengchao Hu, Li Shen, Ya Zhang, Dacheng TaoICLR 2024 · 65 citations
Builds on4
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu et al.NeurIPS 2021 · 121 citations
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li et al.AAAI 2020 · 113 citations
- Multi-Agent Determinantal Q-LearningYaodong Yang, Ying Wen, Jun Wang, Liheng Chen et al.ICML 2020 · 83 citations
- Learning in Nonzero-Sum Stochastic Games with PotentialsDavid Henry Mguni, Yutong Wu, Yali Du, Yaodong Yang et al.ICML 2021 · 51 citations
Related papers
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang et al.NeurIPS 2021 · 73 citations
- Order Matters: Agent-by-agent Policy OptimizationXihuai Wang, Zheng Tian, Ziyu Wan, Ying Wen et al.ICLR 2023 · 3 citations
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement LearningZejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong et al.AAAI 2026
- Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceWeiye Zhao, Feihan Li, Yifan Sun, Rui Chen et al.ICML 2024 · 5 citations
