Prioritized Generative Replay
Renhao Wang, Kevin Frans, Pieter Abbeel, Sergey Levine, Alexei A. Efros
Abstract
Sample-efficient online reinforcement learning often uses replay buffers to store experience for reuse when updating the value function. However, uniform replay is inefficient, since certain classes of transitions can be more relevant to learning. While prioritization of more useful samples is helpful, this strategy can also lead to overfitting, as useful samples are likely to be more rare. In this work, we instead propose a prioritized, parametric version of an agent's memory, using generative models to capture online experience. This paradigm enables (1) densification of past experience, with new generations that benefit from the generative model's generalization capacity and (2) guidance via a family of "relevance functions" that push these generations towards more useful parts of an agent's acquired history. We show this recipe can be instantiated using conditional diffusion models and simple relevance functions such as curiosity-or value-based metrics. Our approach consistently improves performance and sample efficiency in both state-and pixelbased domains. We expose the mechanisms underlying these gains, showing how guidance promotes diversity in our generated transitions and reduces overfitting. We also showcase how our approach can train policies with even higher update-todata ratios than before, opening up avenues to better scale online RL agents. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0ad691e-9566-4875-b349-fd6cdb3d01d5Cited by top-tier papers8
- State-Covering Trajectory Stitching for Diffusion PlannersKyowoon Lee, Jaesik ChoiNeurIPS 2025 · 17 citations
- BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement LearningYunpeng Qing, Yixiao Chi, Shuo Chen, Shunyu Liu et al.ICML 2026 · 4 citations
- Exploratory Diffusion Model for Unsupervised Reinforcement LearningChengyang Ying, Huayu Chen, Xinning Zhou, Zhongkai Hao et al.ICLR 2026 · 4 citations
- Analytic Energy-Guided Policy Optimization for Offline Reinforcement LearningJifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu et al.NeurIPS 2025 · 4 citations
- Off-policy Reinforcement Learning with Model-based Exploration AugmentationLikun Wang, Xiangteng Zhang, Yinuo Wang, Guojian Zhan et al.NeurIPS 2025 · 3 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
Related papers
- The Benefits of Model-Based Generalization in Reinforcement LearningKenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen SchmidhuberICML 2023 · 18 citations
- Synthetic Experience ReplayCong Lu, Philip J. Ball, Yee Whye Teh, Jack Parker-HolderNeurIPS 2023 · 148 citations
- ATraDiff: Accelerating Online Reinforcement Learning with Imaginary TrajectoriesQianlan Yang, Yu-Xiong WangICML 2024 · 2 citations
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 73 citations
- Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement LearningXu-Hui Liu, Tian-Shuo Liu, Shengyi Jiang, Ruifeng Chen et al.ICML 2024 · 10 citations
