Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
Qian Long, Zihan Zhou, Abhinav Gupta, Fei Fang, Yi Wu, Xiaolong Wang
摘要
In fully cooperative environments, agents aim to learn a joint policy to achieve a shared goal. However, existing Multi-Agent Reinforcement Learning (MARL) approaches struggle when scaling to complex coordination tasks. However, as the complexity of joint tasks increases and the policy space expands, agents face significant challenges in achieving the convergence of optimal policies. The limited observational capabilities of agents, coupled with time-varying interaction weights among neighboring agents, lead to challenges in maintaining stable policy evaluations. To address these challenges, we propose GDE, a MARL framework that combines Graph-based value Decomposition with staged Evolutionary policy optimization. To enhance the efficiency of policy exploration and convergence, we use Evolutionary Algorithms (EAs) with diverse in-population characteristics to conduct gradient-free random search. We employ Graph Neural Networks (GNNs) to extend agents' receptive fields, improving information propagation across neighbors and enhancing coordination in dynamic environments without requiring state consensus. Furthermore, the permutation invariance of topological graphs allows GNNs to maintain stable convergence when processing dynamic data. The formation of multiple agent teams enhances GNNs' ability to capture complex coordination dynamics within the multi-agent system. Our method enables staged optimization of agent policies through evolutionary mechanisms while continuously updating joint policies based on graph relationships. Experiments conducted on micro-management in StarCraft II, robot cooperation in MAMuJoCo, and autonomous driving in SUMO demonstrate the superior performance of GDE, validating the effectiveness and necessity of each proposed module. Our code is available: https://github.com/MercyM/GDE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Shared Experience Actor-Critic for Multi-Agent Reinforcement LearningFilippos Christianos, Lukas Schäfer, Stefano V. AlbrechtNeurIPS 2020 · 被引用 238 次
- Randomized Entity-wise Factorization for Multi-Agent Reinforcement LearningShariq Iqbal, Christian A. Schröder de Witt, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 84 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Coach-Player Multi-agent Reinforcement Learning for Dynamic Team CompositionBo Liu, Qiang Liu, Peter Stone, Animesh Garg 等ICML 2021 · 被引用 64 次
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu 等ICLR 2021 · 被引用 63 次
它引用的顶会 Paper1
相关 Paper
- GRDC: A Unified Graph-Driven Framework for Role Discovery and Communication in Multi-Agent Reinforcement LearningZihong Gao, Hongjian Liang, Yuanhui Hao, Lei Hao 等AAAI 2026
- Graph-Supported Dynamic Algorithm Configuration for Multi-Objective Combinatorial OptimizationRobbert Reijnen, Yaoxin Wu, Zaharah Bukhsh, Yingqian ZhangICML 2025
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationSomdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer 等ICML 2020 · 被引用 70 次
- Graph Diffusion for Robust Multi-Agent CoordinationXianghua Zeng, Hang Su, Zhengyi Wang, Zhiyuan LinICML 2025
- RACE: Improve Multi-Agent Reinforcement Learning with Representation Asymmetry and Collaborative EvolutionPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng 等ICML 2023 · 被引用 31 次
