Goal-Conditioned Predictive Coding for Offline Reinforcement Learning
Zilai Zeng, Ce Zhang, Shijie Wang, Chen Sun
Abstract
Recent work has demonstrated the effectiveness of formulating decision making as supervised learning on offline-collected trajectories. Powerful sequence models, such as GPT or BERT, are often employed to encode the trajectories. However, the benefits of performing sequence modeling on trajectory data remain unclear. In this work, we investigate whether sequence modeling has the ability to condense trajectories into useful representations that enhance policy learning. We adopt a two-stage framework that first leverages sequence models to encode trajectory-level representations, and then learns a goal-conditioned policy employing the encoded representations as its input. This formulation allows us to consider many existing supervised offline RL methods as specific instances of our framework. Within this framework, we introduce Goal-Conditioned Predictive Coding (GCPC), a sequence modeling objective that yields powerful trajectory representations and leads to performant policies. Through extensive empirical evaluations on AntMaze, FrankaKitchen and Locomotion environments, we observe that sequence modeling can have a significant impact on challenging decision making tasks. Furthermore, we demonstrate that GCPC learns a goal-conditioned latent representation encoding the future trajectory, which enables competitive performance on all three benchmarks. Our code is available at https://brown-palm.github.io/GCPC/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36ea88ea-2a35-4d4a-9672-61a1b01457d9Cited by top-tier papers8
- Flattening Hierarchies with Policy BootstrappingJohn L. Zhou, Jonathan C. KaoNeurIPS 2025 · 8 citations
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-MakingFan Feng, Selena Ge, Minghao Fu, Zijian Li et al.ICLR 2026 · 3 citations
- A Driving-Style-Adaptive Framework for Vehicle Trajectory PredictionDi Wen, Yu Wang, Zhigang Wu, Zhaocheng He et al.NeurIPS 2025 · 1 citation
- OGBench: Benchmarking Offline Goal-Conditioned RLSeohong Park, Kevin Frans, Benjamin Eysenbach, Sergey LevineICLR 2025
- VideoWorld: Exploring Knowledge Learning from Unlabeled VideosZhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao et al.CVPR 2025
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
Related papers
- M^3PC: Test-time Model Predictive Control using Pretrained Masked Trajectory ModelKehan Wen, Yutong Hu, Yao Mu, Lei KeICLR 2025
- Are Expressive Models Truly Necessary for Offline RL?Guan Wang, Haoyi Niu, Jianxiong Li, Li Jiang et al.AAAI 2025 · 9 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
- Diffused Task-Agnostic Milestone PlannerMineui Hong, Minjae Kang, Songhwai OhNeurIPS 2023 · 15 citations
