Large Batch Experience Replay
Thibault Lahire, Matthieu Geist, Emmanuel Rachelson
摘要
Several algorithms have been proposed to sample non-uniformly the replay buffer of deep Reinforcement Learning (RL) agents to speed-up learning, but very few theoretical foundations of these sampling schemes have been provided. Among others, Prioritized Experience Replay appears as a hyperparameter sensitive heuristic, even though it can provide good performance. In this work, we cast the replay buffer sampling problem as an importance sampling one for estimating the gradient. This allows deriving the theoretically optimal sampling distribution, yielding the best theoretical convergence speed. Elaborating on the knowledge of the ideal sampling scheme, we exhibit new theoretical foundations of Prioritized Experience Replay. The optimal sampling distribution being intractable, we make several approximations providing good results in practice and introduce, among others, LaBER (Large Batch Experience Replay), an easy-to-code and efficient method for sampling the replay buffer. LaBER, which can be combined with Deep Q-Networks, distributional RL agents or actor-critic methods, yields improved performance over a diverse range of Atari games and PyBullet environments, compared to the base agent it is implemented on and to other prioritization schemes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Small batch deep reinforcement learningJohan S. Obando-Ceron, Marc G. Bellemare, Pablo Samuel CastroNeurIPS 2023 · 被引用 38 次
- Prioritizing Samples in Reinforcement Learning with Reducible LossShivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi KahouNeurIPS 2023 · 被引用 36 次
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 被引用 20 次
- Analysis of Stochastic Processes through Replay BuffersShirli Di-Castro Shashua, Shie Mannor, Dotan Di CastroICML 2022 · 被引用 9 次
- Rebalancing Return Coverage for Conditional Sequence Modeling in Offline Reinforcement LearningWensong Bai, Chufan Chen, Yichao Fu, Qihang Xu 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper3
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 被引用 125 次
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 被引用 85 次
相关 Paper
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 被引用 3 次
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 被引用 71 次
- Variance Reduction via Resampling and Experience ReplayJiale Han, Xiaowu Dai, Yuhua ZhuAAAI 2026
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White 等AAAI 2022 · 被引用 14 次
- Regret Minimization Experience Replay in Off-Policy Reinforcement LearningXu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang 等NeurIPS 2021 · 被引用 51 次
