Multimodal Dual Population Evolutionary Reinforcement Learning
Yao Zhang, Ping Huang, Rui Zhang
Abstract
The integration of Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) offers a new paradigm for complex decision-making tasks, especially through methods that optimize actor populations and critics collaboratively (e.g., DBCEM-TD3, QD-RL), collectively referred to as the Actor-Population Critic (APC) framework. However, existing methods still face two major challenges: traditional feature inputs fail to fully capture the spatial relationships, and the design of the critic struggles to balance both the quality and diversity of the policies. To address these issues, we propose Multimodal Dual Population Evolutionary Reinforcement Learning (M-DPERL), which achieves breakthroughs through cross-modal feature augmentation and dual-population coevolution. On one hand, a feature-image bimodal input enhancement mechanism is proposed, which dynamically encodes environmental features into spatial heatmaps. On the other hand, this method introduces a critic population into APC and, for the first time, proposes a population-guided fitness metric to optimize the critic's ability to guide the actor population in balancing quality and diversity. Additionally, we design the Flow-Fix Dynamics (FFD) mechanism to regulate the update rhythm of the dual populations and alleviate the evolutionary chaos in their coevolution. The results across a series of MUJOCO tasks demonstrate that M-DPERL significantly outperforms the baselines, with a 19.2% improvement in sample efficiency and a 17.1% increase in final performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e77e0094-5fe2-4d8a-a1e0-9e6042b6e753Related papers
- Cooperative Heterogeneous Deep Reinforcement LearningHan Zheng, Pengfei Wei, Jing Jiang, Guodong Long et al.NeurIPS 2020 · 20 citations
- Double Buffers CEM-TD3: More Efficient Evolution and Richer ExplorationSheng Zhu, Chun Shen, Shuai Lü, Junhong Wu et al.AAAI 2024 · 3 citations
- Sample-Efficient Quality-Diversity by Cooperative CoevolutionKe Xue, Ren-Jian Wang, Pengyi Li, Dong Li et al.ICLR 2024 · 17 citations
- Two-Stage Evolutionary Reinforcement Learning for Enhancing Exploration and ExploitationQingling Zhu, Xiaoqiang Wu, Qiuzhen Lin, Wei-Neng ChenAAAI 2024 · 9 citations
- ERL-TD: Evolutionary Reinforcement Learning Enhanced with Truncated Variance and Distillation MutationQiuzhen Lin, Yangfan Chen, Lijia Ma, Wei-Neng Chen et al.AAAI 2024 · 6 citations
