Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
Ge Li, Hongyi Zhou, Dominik Roth, Serge Thilges, Fabian Otto, Rudolf Lioutikov, Gerhard Neumann
摘要
Current advancements in reinforcement learning (RL) have predominantly focused on learning step-based policies that generate actions for each perceived state. While these methods efficiently leverage step information from environmental interaction, they often ignore the temporal correlation between actions, resulting in inefficient exploration and unsmooth trajectories that are challenging to implement on real hardware. Episodic RL (ERL) seeks to overcome these challenges by exploring in parameters space that capture the correlation of actions. However, these approaches typically compromise data efficiency, as they treat trajectories as opaque black boxes. In this work, we introduce a novel ERL algorithm, Temporally-Correlated Episodic RL (TCE), which effectively utilizes step information in episodic policy updates, opening the 'black box' in existing ERL methods while retaining the smooth and consistent exploration in parameter space. TCE synergistically combines the advantages of step-based and episodic RL, achieving comparable performance to recent ERL methods while maintaining data efficiency akin to state-of-the-art (SoTA) step-based RL. Our code is available at https://github.com/BruceGeLi/TCE_RL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningHongyi Zhou, Weiran Liao, Xi Huang, Yucheng Tang 等NeurIPS 2025 · 被引用 32 次
- Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of ExpertsOnur Celik, Aleksandar Taranovic, Gerhard NeumannICML 2024 · 被引用 19 次
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 被引用 18 次
- TROLL: Trust Regions Improve Reinforcement Learning for Large Language ModelsPhilipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto 等ICLR 2026 · 被引用 8 次
- Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free ControlDevdhar Patel, Hava T. SiegelmannICLR 2025
它引用的顶会 Paper8
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Neural Dynamic Policies for End-to-End Sensorimotor LearningShikhar Bahl, Mustafa Mukadam, Abhinav Gupta, Deepak PathakNeurIPS 2020 · 被引用 97 次
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 被引用 41 次
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 被引用 38 次
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 被引用 30 次
相关 Paper
- TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement LearningGe Li, Dong Tian, Hongyi Zhou, Xinkai Jiang 等ICLR 2025
- Episodic Reinforcement Learning with Associative MemoryGuangxiang Zhu, Zichuan Lin, Guangwen Yang, Chongjie ZhangICLR 2020 · 被引用 56 次
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie 等ICLR 2023 · 被引用 5 次
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 等NeurIPS 2021 · 被引用 25 次
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 被引用 24 次
