Model-augmented Prioritized Experience Replay
Youngmin Oh, Jinwoo Shin, Eunho Yang, Sung Ju Hwang
Abstract
Experience replay is an essential component in off-policy model-free reinforcement learning (MfRL). Due to its effectiveness, various methods for calculating priority scores on experiences have been proposed for sampling. Since critic networks are crucial to policy learning, TD-error, directly correlated to Q-values, is one of the most frequently used features to compute the scores. However, critic networks often under-or overestimate Q-values, so it is often ineffective to learn for predicting Q-values by sampled experiences based heavily on TD-error. Accordingly, it is valuable to find auxiliary features, which positively support TD-error in calculating the scores for efficient sampling. Motivated by this, we propose a novel experience replay method, which we call model-augmented prioritized experience replay (MaPER), that employs new learnable features driven from components in modelbased RL (MbRL) to calculate the scores on experiences. The proposed MaPER brings the effect of curriculum learning for predicting Q-values better by the critic network with negligible memory and computational overhead compared to the vanilla PER. Indeed, our experimental results on various tasks demonstrate that MaPER can significantly improve the performance of the state-of-the-art offpolicy MfRL and MbRL which includes off-policy MfRL algorithms in its policy optimization procedure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04ab3c02-bb6f-437f-9904-f346e6f37f85Cited by top-tier papers8
- Prioritizing Samples in Reinforcement Learning with Reducible LossShivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi KahouNeurIPS 2023 · 36 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
- Curious Replay for Model-based AdaptationIsaac Kauvar, Chris Doyle, Linqi Zhou, Nick HaberICML 2023 · 18 citations
- Uncertainty-Guided Exploration for Efficient AlphaZero TrainingScott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. KandemirNeurIPS 2025
- Replay Memory as An Empirical MDP: Combining Conservative Estimation with Experience ReplayHongming Zhang, Chenjun Xiao, Han Wang, Jun Jin et al.ICLR 2023
Builds on18
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 389 citations
Related papers
- Prioritized Model Experience ReplayMuxi Tao, jiangtao wen, Yuxing HanICML 2026
- Regret Minimization Experience Replay in Off-Policy Reinforcement LearningXu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang et al.NeurIPS 2021 · 51 citations
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 3 citations
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 85 citations
- Learning to Sample with Local and Global Contexts in Experience Replay BufferYoungmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang et al.ICLR 2021 · 19 citations
