Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement Learning
Julien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai, Chris Pal
Abstract
In multi-agent reinforcement learning, discovering successful collective behaviors is challenging as it requires exploring a joint action space that grows exponentially with the number of agents. While the tractability of independent agent-wise exploration is appealing, this approach fails on tasks that require elaborate group strategies. We argue that coordinating the agents' policies can guide their exploration and we investigate techniques to promote such an inductive bias. We propose two policy regularization methods: TeamReg, which is based on interagent action predictability and CoachReg that relies on synchronized behavior selection. We evaluate each approach on four challenging continuous control tasks with sparse rewards that require varying levels of coordination as well as on the discrete action Google Research Football environment. Our experiments show improved performance across many cooperative multi-agent problems. Finally, we analyze the effects of our proposed methods on the policies that our agents learn and show that our methods successfully enforce the qualities that we propose as proxies for coordinated behaviors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao et al.NeurIPS 2021 · 224 citations
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- Adaptive Rational Activations to Boost Deep Reinforcement LearningQuentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina et al.ICLR 2024 · 25 citations
- ELIGN: Expectation Alignment as a Multi-Agent Intrinsic RewardZixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein et al.NeurIPS 2022 · 22 citations
- Learning to Guide and to be Guided in the Architect-Builder ProblemPaul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier et al.ICLR 2022 · 5 citations
Builds on1
Related papers
- Individual Reward Assisted Multi-Agent Reinforcement LearningLi Wang, Yupeng Zhang, Yujing Hu, Weixun Wang et al.ICML 2022 · 41 citations
- LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningMingyu Yang, Jian Zhao, Xunhan Hu, Wengang Zhou et al.NeurIPS 2022 · 61 citations
- Neighborhood Cognition Consistent Multi-Agent Reinforcement LearningHangyu Mao, Wulong Liu, Jianye Hao, Jun Luo et al.AAAI 2020 · 85 citations
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu et al.NeurIPS 2023 · 37 citations
- Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement LearningRui Tang, Biao Luo, Yongzheng CuiAAAI 2026
