Off-Policy Actor-Critic with Shared Experience Replay
Simon Schmitt, Matteo Hessel, Karen Simonyan
摘要
We investigate the combination of actor-critic reinforcement learning algorithms with uniform large-scale experience replay and propose solutions for two challenges: (a) efficient actor-critic learning with experience replay (b) stability of off-policy learning where agents learn from other agents behaviour. We employ those insights to accelerate hyper-parameter sweeps in which all participating agents run concurrently and share their experience via a common replay module. To this end we analyze the bias-variance tradeoffs in V-trace, a form of importance sampling for actor-critic methods. Based on our analysis, we then argue for mixing experience sampled from replay with on-policy experience, and propose a new trust region scheme that scales effectively to data distributions where V-trace becomes unstable. We provide extensive empirical validation of the proposed solution. We further show the benefits of this setup by demonstrating state-of-the-art data efficiency on Atari among agents trained up until 200M environment frames.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme 等ICML 2020 · 被引用 131 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez 等ICML 2021 · 被引用 69 次
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt 等ICLR 2022 · 被引用 62 次
它引用的顶会 Paper1
相关 Paper
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White 等AAAI 2022 · 被引用 14 次
- Shared Experience Actor-Critic for Multi-Agent Reinforcement LearningFilippos Christianos, Lukas Schäfer, Stefano V. AlbrechtNeurIPS 2020 · 被引用 238 次
- Finite-Sample Analysis of Off-Policy Natural Actor-Critic AlgorithmSajad Khodadadian, Zaiwei Chen, Siva Theja MaguluriICML 2021 · 被引用 33 次
- Efficient Off-Policy Learning for High-Dimensional Action SpacesFabian Otto, Philipp Becker, Ngo Anh Vien, Gerhard NeumannICLR 2025
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 被引用 18 次
