CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning
Jiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura, Hongyuan Zha
摘要
A variety of cooperative multi-agent control problems require agents to achieve individual goals while contributing to collective success. This multi-goal multi-agent setting poses difficulties for recent algorithms, which primarily target settings with a single global reward, due to two new challenges: efficient exploration for learning both individual goal attainment and cooperation for others' success, and credit-assignment for interactions between actions and goals of different agents. To address both challenges, we restructure the problem into a novel two-stage curriculum, in which single-agent goal attainment is learned prior to learning multi-agent cooperation, and we derive a new multi-goal multi-agent policy gradient with a credit function for localized credit assignment. We use a function augmentation scheme to bridge value and policy functions across the curriculum. The complete architecture, called CM3, learns significantly faster than direct adaptations of existing algorithms on three challenging multi-goal multi-agent problems: cooperative navigation in difficult formations, negotiating multi-vehicle lane changes in the SUMO traffic simulator, and strategic cooperation in a Checkers environment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ALMA: Hierarchical Learning for Composite Multi-Agent TasksShariq Iqbal, Robby Costales, Fei ShaNeurIPS 2022 · 被引用 40 次
- Asynchronous Actor-Critic for Multi-Agent Reinforcement LearningYuchen Xiao, Weihao Tan, Christopher AmatoNeurIPS 2022 · 被引用 35 次
- Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized TeamingSachin G. Konan, Esmaeil Seraj, Matthew C. GombolayICLR 2022 · 被引用 27 次
- Mamba: Bringing Multi-Dimensional ABR to WebRTCYueheng Li, Zicheng Zhang, Hao Chen, Zhan MaACM MM 2023 · 被引用 15 次
- Recursive Reasoning Graph for Multi-Agent Reinforcement LearningXiaobai Ma, David Isele, Jayesh K. Gupta, Kikuo Fujimura 等AAAI 2022 · 被引用 8 次
相关 Paper
- Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy GradientWubing Chen, Wenbin Li, Xiao Liu, Shangdong Yang 等AAAI 2023 · 被引用 11 次
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 被引用 159 次
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent ProblemsJiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma 等NeurIPS 2021 · 被引用 48 次
- Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement LearningNicolas Castanet, Olivier Sigaud, Sylvain LamprierICML 2023 · 被引用 6 次
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang 等ICML 2020 · 被引用 64 次
