Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy Gradient
Wubing Chen, Wenbin Li, Xiao Liu, Shangdong Yang, Yang Gao
Abstract
Cooperative multi-agent policy gradient (MAPG) algorithms have recently attracted wide attention and are regarded as a general scheme for the multi-agent system. Credit assignment plays an important role in MAPG and can induce cooperation among multiple agents. However, most MAPG algorithms cannot achieve good credit assignment because of the game-theoretic pathology known as centralized-decentralized mismatch. To address this issue, this paper presents a novel method, Multi-Agent Polarization Policy Gradient (MAPPG). MAPPG takes a simple but efficient polarization function to transform the optimal consistency of joint and individual actions into easily realized constraints, thus enabling efficient credit assignment in MAPPG. Theoretically, we prove that individual policies of MAPPG can converge to the global optimum. Empirically, we evaluate MAPPG on the well-known matrix game and differential game, and verify that MAPPG can converge to the global optimum for both discrete and continuous action spaces. We also evaluate MAPPG on a set of StarCraft II micromanagement tasks and demonstrate that MAPPG outperforms the state-of-the-art MAPG algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40669d3b-a01f-432d-b7c4-9a26fc2863b3Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
Related papers
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang et al.NeurIPS 2021 · 73 citations
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang et al.ICML 2020 · 64 citations
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie et al.ICML 2021 · 88 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
- TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy GradientXingzhou Lou, Junge Zhang, Timothy J. Norman, Kaiqi Huang et al.AAAI 2024 · 2 citations
