Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
Qian Long, Zihan Zhou, Abhinav Gupta, Fei Fang, Yi Wu, Xiaolong Wang
Abstract
In fully cooperative environments, agents aim to learn a joint policy to achieve a shared goal. However, existing Multi-Agent Reinforcement Learning (MARL) approaches struggle when scaling to complex coordination tasks. However, as the complexity of joint tasks increases and the policy space expands, agents face significant challenges in achieving the convergence of optimal policies. The limited observational capabilities of agents, coupled with time-varying interaction weights among neighboring agents, lead to challenges in maintaining stable policy evaluations. To address these challenges, we propose GDE, a MARL framework that combines Graph-based value Decomposition with staged Evolutionary policy optimization. To enhance the efficiency of policy exploration and convergence, we use Evolutionary Algorithms (EAs) with diverse in-population characteristics to conduct gradient-free random search. We employ Graph Neural Networks (GNNs) to extend agents' receptive fields, improving information propagation across neighbors and enhancing coordination in dynamic environments without requiring state consensus. Furthermore, the permutation invariance of topological graphs allows GNNs to maintain stable convergence when processing dynamic data. The formation of multiple agent teams enhances GNNs' ability to capture complex coordination dynamics within the multi-agent system. Our method enables staged optimization of agent policies through evolutionary mechanisms while continuously updating joint policies based on graph relationships. Experiments conducted on micro-management in StarCraft II, robot cooperation in MAMuJoCo, and autonomous driving in SUMO demonstrate the superior performance of GDE, validating the effectiveness and necessity of each proposed module. Our code is available: https://github.com/MercyM/GDE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Shared Experience Actor-Critic for Multi-Agent Reinforcement LearningFilippos Christianos, Lukas Schäfer, Stefano V. AlbrechtNeurIPS 2020 · 238 citations
- Randomized Entity-wise Factorization for Multi-Agent Reinforcement LearningShariq Iqbal, Christian A. Schröder de Witt, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 84 citations
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- Coach-Player Multi-agent Reinforcement Learning for Dynamic Team CompositionBo Liu, Qiang Liu, Peter Stone, Animesh Garg et al.ICML 2021 · 64 citations
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu et al.ICLR 2021 · 63 citations
Builds on1
Related papers
- GRDC: A Unified Graph-Driven Framework for Role Discovery and Communication in Multi-Agent Reinforcement LearningZihong Gao, Hongjian Liang, Yuanhui Hao, Lei Hao et al.AAAI 2026
- Graph-Supported Dynamic Algorithm Configuration for Multi-Objective Combinatorial OptimizationRobbert Reijnen, Yaoxin Wu, Zaharah Bukhsh, Yingqian ZhangICML 2025
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationSomdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer et al.ICML 2020 · 70 citations
- Graph Diffusion for Robust Multi-Agent CoordinationXianghua Zeng, Hang Su, Zhengyi Wang, Zhiyuan LinICML 2025
- RACE: Improve Multi-Agent Reinforcement Learning with Representation Asymmetry and Collaborative EvolutionPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng et al.ICML 2023 · 31 citations
