Future-conditioned Unsupervised Pretraining for Decision Transformer
Zhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu, Wei Yang, Shuai Li
摘要
Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promising, return conditioning is limited to training data labeled with rewards and therefore faces challenges in learning from unsupervised data. In this work, we aim to utilize generalized future conditioning to enable efficient unsupervised pretraining from reward-free and sub-optimal offline data. We propose Pretrained Decision Transformer (PDT), a conceptually simple approach for unsupervised RL pretraining. PDT leverages future trajectory information as a privileged context to predict actions during training. The ability to make decisions based on both present and future factors enhances PDT's capability for generalization. Besides, this feature can be easily incorporated into a return-conditioned framework for online finetuning, by assigning return values to possible futures and sampling future embeddings based on their respective values. Empirically, PDT outperforms or performs on par with its supervised pretraining counterpart, especially when dealing with sub-optimal data. Further analysis reveals that PDT can extract diverse behaviors from offline data and controllably sample high-return behaviors by online finetuning. Code is available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma 等ICLR 2024 · 被引用 43 次
- Unleashing the Power of Pre-trained Language Models for Offline Reinforcement LearningRuizhe Shi, Yuyao Liu, Yanjie Ze, Simon Shaolei Du 等ICLR 2024 · 被引用 36 次
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu 等NeurIPS 2024 · 被引用 31 次
- Reinformer: Max-Return Sequence Modeling for Offline RLZifeng Zhuang, Dengyun Peng, Jinxin Liu, Ziqi Zhang 等ICML 2024 · 被引用 29 次
- Pre-Training Goal-based Models for Sample-Efficient Reinforcement LearningHaoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing LuICLR 2024 · 被引用 26 次
它引用的顶会 Paper23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
相关 Paper
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 被引用 126 次
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
- Prompt Tuning Decision Transformers with Structured and Scalable BanditsFinn Rietz, Oleg Smirnov, Sara Karimi, Lele CaoNeurIPS 2025 · 被引用 2 次
- Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision TransformersKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2024 · 被引用 11 次
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised PretrainingLicong Lin, Yu Bai, Song MeiICLR 2024 · 被引用 74 次
