Generalizable Episodic Memory for Deep Reinforcement Learning
Hao Hu, Jianing Ye, Guangxiang Zhu, Zhizhou Ren, Chongjie Zhang
摘要
Episodic memory-based methods can rapidly latch onto past successful strategies by a non-parametric memory and improve sample efficiency of traditional reinforcement learning. However, little effort is put into the continuous domain, where a state is never visited twice, and previous episodic methods fail to efficiently aggregate experience across trajectories. To address this problem, we propose Generalizable Episodic Memory (GEM), which effectively organizes the state-action values of episodic memory in a generalizable manner and supports implicit planning on memorized trajectories. GEM utilizes a double estimator to reduce the overestimation bias induced by value propagation in the planning process. Empirical evaluation shows that our method significantly outperforms existing trajectory-based methods on various MuJoCo continuous control tasks. To further show the general applicability, we evaluate our method on Atari games with discrete action space, which also shows a significant improvement over baseline algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Offline Reinforcement Learning with Value-based Episodic MemoryXiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang 等ICLR 2022 · 被引用 51 次
- On the Estimation Bias in Double Q-LearningZhizhou Ren, Guangxiang Zhu, Hao Hu, Beining Han 等NeurIPS 2021 · 被引用 35 次
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 等NeurIPS 2021 · 被引用 25 次
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 被引用 20 次
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 被引用 13 次
它引用的顶会 Paper4
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
- Episodic Reinforcement Learning with Associative MemoryGuangxiang Zhu, Zichuan Lin, Guangwen Yang, Chongjie ZhangICLR 2020 · 被引用 56 次
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 被引用 26 次
相关 Paper
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie 等ICLR 2023 · 被引用 5 次
- Memory Based Trajectory-conditioned Policies for Learning from Sparse RewardsYijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng 等NeurIPS 2020 · 被引用 36 次
- CarM: hierarchical episodic memory for continual learningSoobee Lee, Minindu Weerakoon, Jonghyun Choi, Minjia Zhang 等DAC 2022 · 被引用 27 次
- Bridging Imagination and Reality for Model-Based Deep Reinforcement LearningGuangxiang Zhu, Minghao Zhang, Honglak Lee, Chongjie ZhangNeurIPS 2020 · 被引用 24 次
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter 等ICLR 2025
