RL Unplugged: A Collection of Benchmarks for Offline Reinforcement Learning
Çaglar Gülçehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel J. Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li
摘要
Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to learn policies from offline datasets, thus overcoming concerns associated with online data collection in the real-world, including cost, safety, or ethical concerns. In this paper, we propose a benchmark called RL Unplugged to evaluate and compare offline RL methods. RL Unplugged includes data from a diverse range of domains including games (e.g., Atari benchmark) and simulated motor control problems (e.g., DM Control Suite). The datasets include domains that are partially or fully observable, use continuous or discrete actions, and have stochastic vs. deterministic dynamics. We propose detailed evaluation protocols for each domain in RL Unplugged and provide an extensive analysis of supervised learning and offline RL methods using these protocols. We will release data for all our tasks and open-source all algorithms presented in this paper. We hope that our suite of benchmarks will increase the reproducibility of experiments and make it possible to study challenging tasks with a limited computational budget, thus making RL research both more systematic and more accessible across the community. Moving forward, we view RL Unplugged as a living benchmark suite that will evolve and grow with datasets contributed by the research community and ourselves. Our project page is available on github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Why So Pessimistic? Estimating Uncertainties for Offline RL through Ensembles, and Why Their Independence MattersSeyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, Ofir NachumNeurIPS 2022 · 被引用 117 次
- You Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic EnvironmentsKeiran Paster, Sheila A. McIlraith, Jimmy BaNeurIPS 2022 · 被引用 84 次
- Retrieval-Augmented Reinforcement LearningAnirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber 等ICML 2022 · 被引用 69 次
- Data augmentation for efficient learning from parametric expertsAlexandre Galashov, Joshua Scott Merel, Nicolas HeessNeurIPS 2022 · 被引用 9 次
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 被引用 7 次
它引用的顶会 Paper2
相关 Paper
- Benchmarking Offline Reinforcement Learning on Real-Robot HardwareNico Gürtler, Sebastian Blaes, Pavel Kolev, Felix Widmaier 等ICLR 2023 · 被引用 11 次
- Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement LearningBrett Barkley, David Fridovich-KeilICML 2025
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker 等ICLR 2021 · 被引用 112 次
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru 等ICLR 2021 · 被引用 52 次
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
