Goal-Conditioned Predictive Coding for Offline Reinforcement Learning
Zilai Zeng, Ce Zhang, Shijie Wang, Chen Sun
摘要
Recent work has demonstrated the effectiveness of formulating decision making as supervised learning on offline-collected trajectories. Powerful sequence models, such as GPT or BERT, are often employed to encode the trajectories. However, the benefits of performing sequence modeling on trajectory data remain unclear. In this work, we investigate whether sequence modeling has the ability to condense trajectories into useful representations that enhance policy learning. We adopt a two-stage framework that first leverages sequence models to encode trajectory-level representations, and then learns a goal-conditioned policy employing the encoded representations as its input. This formulation allows us to consider many existing supervised offline RL methods as specific instances of our framework. Within this framework, we introduce Goal-Conditioned Predictive Coding (GCPC), a sequence modeling objective that yields powerful trajectory representations and leads to performant policies. Through extensive empirical evaluations on AntMaze, FrankaKitchen and Locomotion environments, we observe that sequence modeling can have a significant impact on challenging decision making tasks. Furthermore, we demonstrate that GCPC learns a goal-conditioned latent representation encoding the future trajectory, which enables competitive performance on all three benchmarks. Our code is available at https://brown-palm.github.io/GCPC/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Flattening Hierarchies with Policy BootstrappingJohn L. Zhou, Jonathan C. KaoNeurIPS 2025 · 被引用 8 次
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-MakingFan Feng, Selena Ge, Minghao Fu, Zijian Li 等ICLR 2026 · 被引用 3 次
- A Driving-Style-Adaptive Framework for Vehicle Trajectory PredictionDi Wen, Yu Wang, Zhigang Wu, Zhaocheng He 等NeurIPS 2025 · 被引用 1 次
- OGBench: Benchmarking Offline Goal-Conditioned RLSeohong Park, Kevin Frans, Benjamin Eysenbach, Sergey LevineICLR 2025
- VideoWorld: Exploring Knowledge Learning from Unlabeled VideosZhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao 等CVPR 2025
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
相关 Paper
- M^3PC: Test-time Model Predictive Control using Pretrained Masked Trajectory ModelKehan Wen, Yutong Hu, Yao Mu, Lei KeICLR 2025
- Are Expressive Models Truly Necessary for Offline RL?Guan Wang, Haoyi Niu, Jianxiong Li, Li Jiang 等AAAI 2025 · 被引用 9 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 被引用 126 次
- Diffused Task-Agnostic Milestone PlannerMineui Hong, Minjae Kang, Songhwai OhNeurIPS 2023 · 被引用 15 次
