Representation Matters: Offline Pretraining for Sequential Decision Making
Mengjiao Yang, Ofir Nachum
摘要
The recent success of supervised learning methods on ever larger offline datasets has spurred interest in the reinforcement learning (RL) field to investigate whether the same paradigms can be translated to RL algorithms. This research area, known as offline RL, has largely focused on offline policy optimization, aiming to find a return-maximizing policy exclusively from offline data. In this paper, we consider a slightly different approach to incorporating offline data into sequential decision-making. We aim to answer the question, what unsupervised objectives applied to offline datasets are able to learn state representations which elevate performance on downstream tasks, whether those downstream tasks be online RL, imitation learning from expert demonstrations, or even offline policy optimization based on the same offline dataset? Through a variety of experiments utilizing standard offline RL datasets, we find that the use of pretraining with unsupervised learning objectives can dramatically improve the performance of policy learning algorithms that otherwise yield mediocre performance on their own. Extensive ablations further provide insights into what components of these unsupervised objectives -- e.g., reward prediction, continuous or discrete representations, pretraining or finetuning -- are most important and in which settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee 等NeurIPS 2022 · 被引用 279 次
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 被引用 138 次
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 被引用 125 次
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 被引用 99 次
- Conservative Data Sharing for Multi-Task Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman 等NeurIPS 2021 · 被引用 94 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
相关 Paper
- Future-conditioned Unsupervised Pretraining for Decision TransformerZhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu 等ICML 2023 · 被引用 32 次
- Behavior Prior Representation learning for Offline Reinforcement LearningHongyu Zang, Xin Li, Jie Yu, Chen Liu 等ICLR 2023 · 被引用 3 次
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 等ICLR 2024 · 被引用 17 次
- Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningLanqing Li, Hai Zhang, Xinyu Zhang, Shatong Zhu 等NeurIPS 2024 · 被引用 24 次
- Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningShentao Yang, Yihao Feng, Shujian Zhang, Mingyuan ZhouICML 2022 · 被引用 14 次
