Two-Stage Evolutionary Reinforcement Learning for Enhancing Exploration and Exploitation
Qingling Zhu, Xiaoqiang Wu, Qiuzhen Lin, Wei-Neng Chen
Abstract
The integration of Evolutionary Algorithm (EA) and Reinforcement Learning (RL) has emerged as a promising approach for tackling some challenges in RL, such as sparse rewards, lack of exploration, and brittle convergence properties. However, existing methods often employ actor networks as individuals of EA, which may constrain their exploratory capabilities, as the entire actor population will stop evolving when the critic network in RL falls into local optima. To alleviate this issue, this paper introduces a Two-stage Evolutionary Reinforcement Learning (TERL) framework that maintains a population containing both actor and critic networks. TERL divides the learning process into two stages. In the initial stage, individuals independently learn actor-critic networks, which are optimized alternatively by RL and Particle Swarm Optimization (PSO). This dual optimization fosters greater exploration, curbing susceptibility to local optima. Shared information from a common replay buffer and PSO algorithm substantially mitigates the computational load of training multiple agents. In the subsequent stage, TERL shifts to a refined exploitation phase. Here, only the best individual undergoes further refinement, while the remaining individuals continue PSO-based optimization. This allocates more computational resources to the best individual for yielding superior performance. Empirical assessments, conducted across a range of continuous control problems, validate the efficacy of the proposed TERL paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec4f378f-7fa3-46a5-9f65-d2b297329703Builds on3
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 101 citations
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationSomdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer et al.ICML 2020 · 70 citations
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng et al.ICLR 2023 · 16 citations
Related papers
- Cooperative Heterogeneous Deep Reinforcement LearningHan Zheng, Pengfei Wei, Jing Jiang, Guodong Long et al.NeurIPS 2020 · 20 citations
- An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite IndividualsYangyang Zhao, Ben Niu, Libo Qin, Shihan WangACL 2025 · 3 citations
- Multimodal Dual Population Evolutionary Reinforcement LearningYao Zhang, Ping Huang, Rui ZhangACM MM 2025
- ERL-TD: Evolutionary Reinforcement Learning Enhanced with Truncated Variance and Distillation MutationQiuzhen Lin, Yangfan Chen, Lijia Ma, Wei-Neng Chen et al.AAAI 2024 · 6 citations
- EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous ControlSamuel Holt, Todor Davchev, Dhruva Tirumala, Ben Moran et al.ICML 2025
