Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement Learning
Julien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai, Chris Pal
摘要
In multi-agent reinforcement learning, discovering successful collective behaviors is challenging as it requires exploring a joint action space that grows exponentially with the number of agents. While the tractability of independent agent-wise exploration is appealing, this approach fails on tasks that require elaborate group strategies. We argue that coordinating the agents' policies can guide their exploration and we investigate techniques to promote such an inductive bias. We propose two policy regularization methods: TeamReg, which is based on interagent action predictability and CoachReg that relies on synchronized behavior selection. We evaluate each approach on four challenging continuous control tasks with sparse rewards that require varying levels of coordination as well as on the discrete action Google Research Football environment. Our experiments show improved performance across many cooperative multi-agent problems. Finally, we analyze the effects of our proposed methods on the policies that our agents learn and show that our methods successfully enforce the qualities that we propose as proxies for coordinated behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao 等NeurIPS 2021 · 被引用 224 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Adaptive Rational Activations to Boost Deep Reinforcement LearningQuentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina 等ICLR 2024 · 被引用 25 次
- ELIGN: Expectation Alignment as a Multi-Agent Intrinsic RewardZixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein 等NeurIPS 2022 · 被引用 22 次
- Learning to Guide and to be Guided in the Architect-Builder ProblemPaul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier 等ICLR 2022 · 被引用 5 次
它引用的顶会 Paper1
相关 Paper
- Individual Reward Assisted Multi-Agent Reinforcement LearningLi Wang, Yupeng Zhang, Yujing Hu, Weixun Wang 等ICML 2022 · 被引用 41 次
- LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningMingyu Yang, Jian Zhao, Xunhan Hu, Wengang Zhou 等NeurIPS 2022 · 被引用 61 次
- Neighborhood Cognition Consistent Multi-Agent Reinforcement LearningHangyu Mao, Wulong Liu, Jianye Hao, Jun Luo 等AAAI 2020 · 被引用 85 次
- Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement LearningYifan Zang, Jinmin He, Kai Li, Haobo Fu 等NeurIPS 2023 · 被引用 37 次
- Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement LearningRui Tang, Biao Luo, Yongzheng CuiAAAI 2026
