Sample-Efficient Reinforcement Learning by Breaking the Replay Ratio Barrier
Pierluca D'Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G. Bellemare, Aaron C. Courville
摘要
Increasing the replay ratio, the number of updates of an agent's parameters per environment interaction, is an appealing strategy for improving the sample efficiency of deep reinforcement learning algorithms. In this work, we show that fully or partially resetting the parameters of deep reinforcement learning agents causes better replay ratio scaling capabilities to emerge. We push the limits of the sample efficiency of carefully-modified algorithms by training them using an order of magnitude more updates than usual, significantly improving their performance in the Atari 100k and DeepMind Control Suite benchmarks. We then provide an analysis of the design choices required for favorable replay ratio scaling to be possible and discuss inherent limits and tradeoffs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper76
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare 等ICML 2023 · 被引用 155 次
- STORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningWeipu Zhang, Gang Wang, Jian Sun, Yetian Yuan 等NeurIPS 2023 · 被引用 154 次
- Synthetic Experience ReplayCong Lu, Philip J. Ball, Yee Whye Teh, Jack Parker-HolderNeurIPS 2023 · 被引用 148 次
相关 Paper
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningHojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak 等NeurIPS 2023 · 被引用 50 次
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- MAD-TD: Model-Augmented Data stabilizes High Update Ratio RLClaas Voelcker, Marcel Hussing, Eric Eaton, Amir-massoud Farahmand 等ICLR 2025
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
