Large Batch Experience Replay
Thibault Lahire, Matthieu Geist, Emmanuel Rachelson
Abstract
Several algorithms have been proposed to sample non-uniformly the replay buffer of deep Reinforcement Learning (RL) agents to speed-up learning, but very few theoretical foundations of these sampling schemes have been provided. Among others, Prioritized Experience Replay appears as a hyperparameter sensitive heuristic, even though it can provide good performance. In this work, we cast the replay buffer sampling problem as an importance sampling one for estimating the gradient. This allows deriving the theoretically optimal sampling distribution, yielding the best theoretical convergence speed. Elaborating on the knowledge of the ideal sampling scheme, we exhibit new theoretical foundations of Prioritized Experience Replay. The optimal sampling distribution being intractable, we make several approximations providing good results in practice and introduce, among others, LaBER (Large Batch Experience Replay), an easy-to-code and efficient method for sampling the replay buffer. LaBER, which can be combined with Deep Q-Networks, distributional RL agents or actor-critic methods, yields improved performance over a diverse range of Atari games and PyBullet environments, compared to the base agent it is implemented on and to other prioritization schemes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5147e544-0aaf-4efa-b559-876fb2546015Cited by top-tier papers6
- Small batch deep reinforcement learningJohan S. Obando-Ceron, Marc G. Bellemare, Pablo Samuel CastroNeurIPS 2023 · 38 citations
- Prioritizing Samples in Reinforcement Learning with Reducible LossShivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi KahouNeurIPS 2023 · 36 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
- Analysis of Stochastic Processes through Replay BuffersShirli Di-Castro Shashua, Shie Mannor, Dotan Di CastroICML 2022 · 9 citations
- Rebalancing Return Coverage for Conditional Sequence Modeling in Offline Reinforcement LearningWensong Bai, Chufan Chen, Yichao Fu, Qihang Xu et al.NeurIPS 2025 · 1 citation
Builds on3
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 125 citations
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 85 citations
Related papers
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 3 citations
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 71 citations
- Variance Reduction via Resampling and Experience ReplayJiale Han, Xiaowu Dai, Yuhua ZhuAAAI 2026
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White et al.AAAI 2022 · 14 citations
- Regret Minimization Experience Replay in Off-Policy Reinforcement LearningXu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang et al.NeurIPS 2021 · 51 citations
