Generalizable Episodic Memory for Deep Reinforcement Learning
Hao Hu, Jianing Ye, Guangxiang Zhu, Zhizhou Ren, Chongjie Zhang
Abstract
Episodic memory-based methods can rapidly latch onto past successful strategies by a non-parametric memory and improve sample efficiency of traditional reinforcement learning. However, little effort is put into the continuous domain, where a state is never visited twice, and previous episodic methods fail to efficiently aggregate experience across trajectories. To address this problem, we propose Generalizable Episodic Memory (GEM), which effectively organizes the state-action values of episodic memory in a generalizable manner and supports implicit planning on memorized trajectories. GEM utilizes a double estimator to reduce the overestimation bias induced by value propagation in the planning process. Empirical evaluation shows that our method significantly outperforms existing trajectory-based methods on various MuJoCo continuous control tasks. To further show the general applicability, we evaluate our method on Atari games with discrete action space, which also shows a significant improvement over baseline algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 566a5f8c-adcf-4d82-97cf-5ed47aa9d673Cited by top-tier papers9
- Offline Reinforcement Learning with Value-based Episodic MemoryXiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang et al.ICLR 2022 · 51 citations
- On the Estimation Bias in Double Q-LearningZhizhou Ren, Guangxiang Zhu, Hao Hu, Beining Han et al.NeurIPS 2021 · 35 citations
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran et al.NeurIPS 2021 · 25 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 13 citations
Builds on4
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Episodic Reinforcement Learning with Associative MemoryGuangxiang Zhu, Zichuan Lin, Guangwen Yang, Chongjie ZhangICLR 2020 · 56 citations
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 26 citations
Related papers
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie et al.ICLR 2023 · 5 citations
- Memory Based Trajectory-conditioned Policies for Learning from Sparse RewardsYijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng et al.NeurIPS 2020 · 36 citations
- CarM: hierarchical episodic memory for continual learningSoobee Lee, Minindu Weerakoon, Jonghyun Choi, Minjia Zhang et al.DAC 2022 · 27 citations
- Bridging Imagination and Reality for Model-Based Deep Reinforcement LearningGuangxiang Zhu, Minghao Zhang, Honglak Lee, Chongjie ZhangNeurIPS 2020 · 24 citations
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter et al.ICLR 2025
