Learning Progress Driven Multi-Agent Curriculum
Wenshuai Zhao, Zhiyuan Li, Joni Pajarinen
摘要
The number of agents can be an effective curriculum variable for controlling the difficulty of multi-agent reinforcement learning (MARL) tasks. Existing work typically uses manually defined curricula such as linear schemes. We identify two potential flaws while applying existing reward-based automatic curriculum learning methods in MARL: (1) The expected episode return used to measure task difficulty has high variance; (2) Credit assignment difficulty can be exacerbated in tasks where increasing the number of agents yields higher returns which is common in many MARL tasks. To address these issues, we propose to control the curriculum by using a TD-error based learning progress measure and by letting the curriculum proceed from an initial context distribution to the final task specific one. Since our approach maintains a distribution over the number of agents and measures learning progress rather than absolute performance, which often increases with the number of agents, we alleviate problem (2). Moreover, the learning progress measure naturally alleviates problem (1) by aggregating returns. In three challenging sparse-reward MARL benchmarks, our approach outperforms state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- From Few to More: Large-Scale Dynamic Multiagent Curriculum LearningWeixun Wang, Tianpei Yang, Yong Liu, Jianye Hao 等AAAI 2020 · 被引用 138 次
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang 等ICML 2022 · 被引用 49 次
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent ProblemsJiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma 等NeurIPS 2021 · 被引用 48 次
- Optimistic Multi-Agent Policy GradientWenshuai Zhao, Yi Zhao, Zhiyuan Li, Juho Kannala 等ICML 2024 · 被引用 7 次
- Improving Environment Novelty Quantification for Effective Unsupervised Environment DesignJayden Teoh, Wenjun Li, Pradeep VarakanthamNeurIPS 2024 · 被引用 6 次
相关 Paper
- PORTAL: Automatic Curricula Generation for Multiagent Reinforcement LearningJizhou Wu, Jianye Hao, Tianpei Yang, Xiaotian Hao 等AAAI 2024 · 被引用 12 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
- Automated curriculum generation through setter-solver interactionsSébastien Racanière, Andrew K. Lampinen, Adam Santoro, David P. Reichert 等ICLR 2020 · 被引用 41 次
- Curriculum Reinforcement Learning via Constrained Optimal TransportPascal Klink, Haoyi Yang, Carlo D'Eramo, Jan Peters 等ICML 2022 · 被引用 44 次
- Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationPeide Huang, Mengdi Xu, Jiacheng Zhu, Laixi Shi 等NeurIPS 2022 · 被引用 44 次
