Mutual-Information Regularized Multi-Agent Policy Iteration
Jiangxing Wang, Deheng Ye, Zongqing Lu
Abstract
Despite the success of cooperative multi-agent reinforcement learning algorithms, most of them focus on a single team composition, which prevents them from being used in more realistic scenarios where dynamic team composition is possible. While some studies attempt to solve this problem via multi-task learning in a fixed set of team compositions, there is still a risk of overfitting to the training set, which may lead to catastrophic performance when facing dramatically varying team compositions during execution. To address this problem, we propose to use mutual information (MI) as an augmented reward to prevent individual policies from relying too much on team-related information and encourage agents to learn policies that are robust in different team compositions. Optimizing this MI-augmented objective in an off-policy manner can be intractable due to the existence of dynamic marginal distribution. To alleviate this problem, we first propose a multi-agent policy iteration algorithm with a fixed marginal distribution and prove its convergence and optimality. Then, we propose to employ the Blahut–Arimoto algorithm and an imaginary team composition distribution for optimization with approximate marginal distribution as the practical implementation. Empirically, our method demonstrates strong zero-shot generalization to dynamic team compositions in complex cooperative tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9711c4d-ca9d-48b5-b503-68343354f1faCited by top-tier papers1
Ask how each one uses itBuilds on21
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen et al.NeurIPS 2020 · 225 citations
Related papers
- Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized TeamingSachin G. Konan, Esmaeil Seraj, Matthew C. GombolayICLR 2022 · 27 citations
- Learning Policy Committees for Effective Personalization in MDPs with Diverse TasksLuise Ge, Michael Lanier, Anindya Sarkar, Bengisu Guresti et al.ICML 2025
- Optimistic Multi-Agent Policy GradientWenshuai Zhao, Yi Zhao, Zhiyuan Li, Juho Kannala et al.ICML 2024 · 7 citations
- Coach-Player Multi-agent Reinforcement Learning for Dynamic Team CompositionBo Liu, Qiang Liu, Peter Stone, Animesh Garg et al.ICML 2021 · 64 citations
- Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement LearningJunkai Zhang, Jinmin He, Yifan Zhang, Yifan Zang et al.NeurIPS 2025 · 1 citation
