ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy Representation
Jianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng, Xian Fu, Zhaopeng Meng
摘要
Deep Reinforcement Learning (Deep RL) and Evolutionary Algorithms (EA) are two major paradigms of policy optimization with distinct learning principles, i.e., gradient-based v.s. gradient-free. An appealing research direction is integrating Deep RL and EA to devise new methods by fusing their complementary advantages. However, existing works on combining Deep RL and EA have two common drawbacks: 1) the RL agent and EA agents learn their policies individually, neglecting efficient sharing of useful common knowledge; 2) parameter-level policy optimization guarantees no semantic level of behavior evolution for the EA side. In this paper, we propose Evolutionary Reinforcement Learning with Two-scale State Representation and Policy Representation (ERL-Re), a novel solution to the aforementioned two drawbacks. The key idea of ERL-Re is two-scale representation: all EA and RL policies share the same nonlinear state representation while maintaining individual linear policy representations. The state representation conveys expressive common features of the environment learned by all the agents collectively; the linear policy representation provides a favorable space for efficient policy optimization, where novel behavior-level crossover and mutation operations can be performed. Moreover, the linear policy representation allows convenient generalization of policy fitness with the help of the Policy-extended Value Function Approximator (PeVFA), further improving the sample efficiency of fitness estimation. The experiments on a range of continuous control tasks show that ERL-Re consistently outperforms advanced baselines and achieves the State Of The Art (SOTA). Our code is available on https://github.com/yeshenpy/ERL-Re2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- RACE: Improve Multi-Agent Reinforcement Learning with Representation Asymmetry and Collaborative EvolutionPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng 等ICML 2023 · 被引用 31 次
- Sample-Efficient Quality-Diversity by Cooperative CoevolutionKe Xue, Ren-Jian Wang, Pengyi Li, Dong Li 等ICLR 2024 · 被引用 17 次
- EvoRainbow: Combining Improvements in Evolutionary Reinforcement Learning for Policy SearchPengyi Li, Yan Zheng, Hongyao Tang, Xian Fu 等ICML 2024 · 被引用 13 次
- Value-Evolutionary-Based Reinforcement LearningPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng 等ICML 2024 · 被引用 10 次
- Two-Stage Evolutionary Reinforcement Learning for Enhancing Exploration and ExploitationQingling Zhu, Xiaoqiang Wu, Qiuzhen Lin, Wei-Neng ChenAAAI 2024 · 被引用 9 次
它引用的顶会 Paper14
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
相关 Paper
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 被引用 24 次
- Cooperative Heterogeneous Deep Reinforcement LearningHan Zheng, Pengfei Wei, Jing Jiang, Guodong Long 等NeurIPS 2020 · 被引用 20 次
- ERL-TD: Evolutionary Reinforcement Learning Enhanced with Truncated Variance and Distillation MutationQiuzhen Lin, Yangfan Chen, Lijia Ma, Wei-Neng Chen 等AAAI 2024 · 被引用 6 次
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 被引用 101 次
- On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared RepresentationsGuojun Xiong, Shufan Wang, Daniel Jiang, Jian LiICLR 2025
