Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, Marc G. Bellemare
摘要
Learning tabula rasa, that is without any previously learned knowledge, is the prevalent workflow in reinforcement learning (RL) research. However, RL systems, when applied to large-scale settings, rarely operate tabula rasa. Such large-scale systems undergo multiple design or algorithmic changes during their development cycle and use ad hoc approaches for incorporating these changes without re-training from scratch, which would have been prohibitively expensive. Additionally, the inefficiency of deep RL typically excludes researchers without access to industrialscale resources from tackling computationally-demanding problems. To address these issues, we present reincarnating RL as an alternative workflow or class of problem settings, where prior computational work (e.g., learned policies) is reused or transferred between design iterations of an RL agent, or from one RL agent to another. As a step towards enabling reincarnating RL from any agent to any other agent, we focus on the specific setting of efficiently transferring an existing sub-optimal policy to a standalone value-based RL agent. We find that existing approaches fail in this setting and propose a simple algorithm to address their limitations. Equipped with this algorithm, we demonstrate reincarnating RL's gains over tabula rasa RL on Atari 2600 games, a challenging locomotion task, and the real-world problem of navigating stratospheric balloons. Overall, this work argues for an alternative approach to RL research, which we believe could significantly improve real-world RL adoption and help democratize it further. Open-sourced code and trained agents at agarwl.github.io/reincarnating_rl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk 等ICLR 2024 · 被引用 311 次
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare 等ICML 2023 · 被引用 155 次
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingKolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi 等ICML 2023 · 被引用 110 次
- Deep Reinforcement Learning with Plasticity InjectionEvgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle 等NeurIPS 2023 · 被引用 98 次
它引用的顶会 Paper15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
相关 Paper
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski 等NeurIPS 2021 · 被引用 152 次
- AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud SystemsHaoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke 等USENIX ATC 2023 · 被引用 95 次
- Transient Non-stationarity and Generalisation in Deep Reinforcement LearningMaximilian Igl, Gregory Farquhar, Jelena Luketina, Wendelin Boehmer 等ICLR 2021 · 被引用 104 次
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 被引用 87 次
- Investigating Multi-task Pretraining and Generalization in Reinforcement LearningAdrien Ali Taïga, Rishabh Agarwal, Jesse Farebrother, Aaron C. Courville 等ICLR 2023
