Double Buffers CEM-TD3: More Efficient Evolution and Richer Exploration
Sheng Zhu, Chun Shen, Shuai Lü, Junhong Wu, Daolong An
摘要
CEM-TD3 is a combination scheme using the simple cross-entropy method (CEM) and Twin Delayed Deep Deterministic policy gradient (TD3), and it achieves a satisfactory trade-off between performance and sample efficiency. However, we find that CEM-TD3 cannot fully address the low efficiency of policy search caused by CEM, and the policy gradient learning introduced by TD3 will weaken the diversity of individuals in the population. In this paper, we propose Double Buffers CEM-TD3 (DBCEM-TD3) that optimizes both CEM and TD3. For CEM, DBCEM-TD3 maintains an actor buffer to store the population required for evolution. In each iteration, it only needs to generate a small number of actors to replace the poor actors in the policy buffer to achieve more efficient evolution. The fitness of individuals in the actor buffer decreases exponentially with time, which can avoid premature convergence of the mean actor. For TD3, DBCEM-TD3 maintains a critic buffer with the same number of critics as the number of actors generated in each iteration, and each critic is trained independently by sampling from the shared replay buffer. In each iteration, each newly generated actor uses different critics to guide learning. This ensures more diverse behaviors among the learned actors, enabling richer experiences to be collected during the evaluation phase. We conduct experimental evaluations on five continuous control tasks provided by OpenAI Gym. DBCEM-TD3 outperforms CEM-TD3, TD3, and other classic off-policy reinforcement learning algorithms in terms of performance and sample efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 被引用 138 次
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 被引用 101 次
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 被引用 69 次
- Population-Guided Parallel Policy Search for Reinforcement LearningWhiyoung Jung, Giseung Park, Youngchul SungICLR 2020 · 被引用 41 次
相关 Paper
- Multimodal Dual Population Evolutionary Reinforcement LearningYao Zhang, Ping Huang, Rui ZhangACM MM 2025
- A Simple Decentralized Cross-Entropy MethodZichen Zhang, Jun Jin, Martin Jägersand, Jun Luo 等NeurIPS 2022 · 被引用 12 次
- ERL-TD: Evolutionary Reinforcement Learning Enhanced with Truncated Variance and Distillation MutationQiuzhen Lin, Yangfan Chen, Lijia Ma, Wei-Neng Chen 等AAAI 2024 · 被引用 6 次
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng 等ICLR 2023 · 被引用 16 次
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 被引用 24 次
