RvS: What is Essential for Offline RL via Supervised Learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, Sergey Levine
Abstract
Recent work has shown that supervised learning alone, without temporal difference (TD) learning, can be remarkably effective for offline RL. When does this hold true, and which algorithmic components are necessary? Through extensive experiments, we boil supervised learning for offline RL down to its essential elements. In every environment suite we consider, simply maximizing likelihood with a two-layer feedforward MLP is competitive with state-of-the-art results of substantially more complex methods based on TD learning or sequence modeling with Transformers. Carefully choosing model capacity (e.g., via regularization or architecture) and choosing which information to condition on (e.g., goals or rewards) are critical for performance. These insights serve as a field guide for practitioners doing Reinforcement Learning via Supervised Learning (which we coin"RvS learning"). They also probe the limits of existing RvS methods, which are comparatively weak on random data, and suggest a number of open problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5167a88-41cd-4abf-bae3-99a88bb5f404Cited by top-tier papers76
- OpenChat: Advancing Open-source Language Models with Mixed-Quality DataGuan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li et al.ICLR 2024 · 328 citations
- OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online EnsemblingYifan Zhang, Qingsong Wen, Xue Wang, Weiqi Chen et al.NeurIPS 2023 · 124 citations
- When does return-conditioned supervised learning work for offline reinforcement learning?David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche et al.NeurIPS 2022 · 107 citations
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 99 citations
- You Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic EnvironmentsKeiran Paster, Sheila A. McIlraith, Jimmy BaNeurIPS 2022 · 84 citations
Builds on15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
Related papers
- Offline Multi-Agent Reinforcement Learning with Knowledge DistillationWei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, Phillip IsolaNeurIPS 2022 · 62 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
- Offline Actor-Critic Reinforcement Learning Scales to Large ModelsJost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang, Oliver Groth et al.ICML 2024 · 37 citations
- Bootstrapped Transformer for Offline Reinforcement LearningKerong Wang, Hanye Zhao, Xufang Luo, Kan Ren et al.NeurIPS 2022 · 54 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
