Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, Yaodong Yang
摘要
Trust region methods rigorously enabled reinforcement learning (RL) agents to learn monotonically improving policies, leading to superior performance on a variety of tasks. Unfortunately, when it comes to multi-agent reinforcement learning (MARL), the property of monotonic improvement may not simply apply; this is because agents, even in cooperative games, could have conflicting directions of policy updates. As a result, achieving a guaranteed improvement on the joint policy where each agent acts individually remains an open challenge. In this paper, we extend the theory of trust region learning to cooperative MARL. Central to our findings are the multi-agent advantage decomposition lemma and the sequential policy update scheme. Based on these, we develop Heterogeneous-Agent Trust Region Policy Optimisation (HATPRO) and Heterogeneous-Agent Proximal Policy Optimisation (HAPPO) algorithms. Unlike many existing MARL algorithms, HATRPO/HAPPO do not need agents to share parameters, nor do they need any restrictive assumptions on decomposibility of the joint value function. Most importantly, we justify in theory the monotonic improvement property of HATRPO/HAPPO. We evaluate the proposed methods on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that HATRPO and HAPPO significantly outperform strong baselines such as IPPO, MAPPO and MADDPG on all tested tasks, thereby establishing a new state of the art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper73
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- Heterogeneous Agent Q-weighted Policy OptimizationBor-Jiun Lin, Chun-Yi LeeICLR 2026 · 被引用 102 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang 等NeurIPS 2022 · 被引用 65 次
- Learning Multi-Agent Communication from Graph Modeling PerspectiveShengchao Hu, Li Shen, Ya Zhang, Dacheng TaoICLR 2024 · 被引用 65 次
它引用的顶会 Paper4
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu 等NeurIPS 2021 · 被引用 121 次
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li 等AAAI 2020 · 被引用 113 次
- Multi-Agent Determinantal Q-LearningYaodong Yang, Ying Wen, Jun Wang, Liheng Chen 等ICML 2020 · 被引用 83 次
- Learning in Nonzero-Sum Stochastic Games with PotentialsDavid Henry Mguni, Yutong Wu, Yali Du, Yaodong Yang 等ICML 2021 · 被引用 51 次
相关 Paper
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
- Order Matters: Agent-by-agent Policy OptimizationXihuai Wang, Zheng Tian, Ziyu Wan, Ying Wen 等ICLR 2023 · 被引用 3 次
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song 等NeurIPS 2024 · 被引用 22 次
- HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement LearningZejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong 等AAAI 2026
- Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceWeiye Zhao, Feihan Li, Yifan Sun, Rui Chen 等ICML 2024 · 被引用 5 次
