Regret Minimization Experience Replay in Off-Policy Reinforcement Learning
Xu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang, Feng Xu, Yang Yu
Abstract
In reinforcement learning, experience replay stores past samples for further reuse. Prioritized sampling is a promising technique to better utilize these samples. Previous criteria of prioritization include TD error, recentness and corrective feedback, which are mostly heuristically designed. In this work, we start from the regret minimization objective, and obtain an optimal prioritization strategy for Bellman update that can directly maximize the return of the policy. The theory suggests that data with higher hindsight TD error, better on-policiness and more accurate Q value should be assigned with higher weights during sampling. Thus most previous criteria only consider this strategy partially. We not only provide theoretical justifications for previous criteria, but also propose two new methods to compute the prioritization weight, namely ReMERN and ReMERT. ReMERN learns an error network, while ReMERT exploits the temporal ordering of states. Both methods outperform previous prioritized sampling algorithms in challenging RL benchmarks, including MuJoCo, Atari and Meta-World.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Attentive Experience ReplayPeiquan Sun, Wengang Zhou, Houqiang LiAAAI 2020 · 62 citations
- Prioritizing Samples in Reinforcement Learning with Reducible LossShivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi KahouNeurIPS 2023 · 36 citations
- State Regularized Policy Optimization on Data with Dynamics ShiftZhenghai Xue, Qingpeng Cai, Shuchang Liu, Dong Zheng et al.NeurIPS 2023 · 30 citations
- Baffle: Hiding Backdoors in Offline Reinforcement Learning DatasetsChen Gong, Zhou Yang, Yunpeng Bai, Junda He et al.S&P 2024 · 28 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
Builds on5
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 124 citations
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 85 citations
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 38 citations
Related papers
- Model-augmented Prioritized Experience ReplayYoungmin Oh, Jinwoo Shin, Eunho Yang, Sung Ju HwangICLR 2022 · 21 citations
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 3 citations
- Prioritized Model Experience ReplayMuxi Tao, jiangtao wen, Yuxing HanICML 2026
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 18 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
