Replay Memory as An Empirical MDP: Combining Conservative Estimation with Experience Replay
Hongming Zhang, Chenjun Xiao, Han Wang, Jun Jin, Bo Xu, Martin Müller
Abstract
Experience replay, which stores transitions in a replay memory for repeated use, plays an important role of improving sample efficiency in reinforcement learning. Existing techniques such as reweighted sampling, episodic learning and reverse sweep update process the information in the replay memory to make experience replay more efficient. In this work, we further exploit the information in the replay memory by treating it as an empirical Replay Memory MDP (RM-MDP). By solving it with dynamic programming, we learn a conservative value estimate that only considers transitions observed in the replay memory. Both value and policy regularizers based on this conservative estimate are developed and integrated with model-free learning algorithms. We design the memory density metric to measure the quality of RM-MDP. Our empirical studies quantitatively find a strong correlation between performance improvement and memory density. Our method combines Conservative Estimation with Experience Replay (CEER), improving sample efficiency by a large margin, especially when the memory density is high. Even when the memory density is low, such a conservative estimate can still help to avoid suicidal actions and thereby improve performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted BehaviorsFengshuo Bai, Runze Liu, Yali Du, Ying Wen et al.AAAI 2025 · 15 citations
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous ManipulationKefei Zhu, Fengshuo Bai, YuanHao Xiang, Yishuai Cai et al.NeurIPS 2025 · 8 citations
- Exploiting the Replay Memory Before Exploring the Environment: Enhancing Reinforcement Learning Through Empirical MDP IterationHongming Zhang, Chenjun Xiao, Chao Gao, Han Wang et al.NeurIPS 2024 · 7 citations
- Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RLYu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang et al.ICML 2024 · 7 citations
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 3 citations
Builds on16
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
Related papers
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li et al.AAAI 2022 · 38 citations
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White et al.AAAI 2022 · 14 citations
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 3 citations
- Episodic Reinforcement Learning with Associative MemoryGuangxiang Zhu, Zichuan Lin, Guangwen Yang, Chongjie ZhangICLR 2020 · 56 citations
- Prioritized Generative ReplayRenhao Wang, Kevin Frans, Pieter Abbeel, Sergey Levine et al.ICLR 2025
