TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy Gradient
Xingzhou Lou, Junge Zhang, Timothy J. Norman, Kaiqi Huang, Yali Du
Abstract
Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. While using individual critics for policy updates can avoid this issue, they severely limit cooperation among agents. To address this issue, we propose an agent topology framework, which decides whether other agents should be considered in policy gradient and achieves compromise between facilitating cooperation and alleviating the CDM issue. The agent topology allows agents to use coalition utility as learning objective instead of global utility by centralized critics or local utility by individual critics. To constitute the agent topology, various models are studied. We propose Topology-based multi-Agent Policy gradiEnt (TAPE) for both stochastic and deterministic MAPG methods. We prove the policy improvement theorem for stochastic TAPE and give a theoretical explanation for the improved cooperation among agents. Experiment results on several benchmarks show the agent topology is able to facilitate agent cooperation and alleviate CDM issue respectively to improve performance of TAPE. Finally, multiple ablation studies and a heuristic graph search algorithm are devised to show the efficacy of the agent topology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40328a3d-1bfc-44ec-af31-fbabfa21347eCited by top-tier papers1
Ask how each one uses itBuilds on9
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao et al.NeurIPS 2021 · 224 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
Related papers
- Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy GradientWubing Chen, Wenbin Li, Xiao Liu, Shangdong Yang et al.AAAI 2023 · 11 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- Optimistic Multi-Agent Policy GradientWenshuai Zhao, Yi Zhao, Zhiyuan Li, Juho Kannala et al.ICML 2024 · 7 citations
- HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement LearningZejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong et al.AAAI 2026
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun et al.ICML 2021 · 66 citations
