Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets
Anirudhan Badrinath, Yannis Flet-Berliac, Allen Nie, Emma Brunskill
Abstract
Despite the recent advancements in offline reinforcement learning via supervised learning (RvS) and the success of the decision transformer (DT) architecture in various domains, DTs have fallen short in several challenging benchmarks. The root cause of this underperformance lies in their inability to seamlessly connect segments of suboptimal trajectories. To overcome this limitation, we present a novel approach to enhance RvS methods by integrating intermediate targets. We introduce the Waypoint Transformer (WT), using an architecture that builds upon the DT framework and conditioned on automatically-generated waypoints. The results show a significant increase in the final return compared to existing RvS methods, with performance on par or greater than existing state-of-the-art temporal difference learning-based methods. Additionally, the performance and stability improvements are largest in the most challenging environments and data configurations, including AntMaze Large Play/Diverse and Kitchen Mixed/Partial.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c266412-ce72-4100-8767-aaa0c3008609Cited by top-tier papers12
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu et al.NeurIPS 2024 · 31 citations
- Pre-Training Goal-based Models for Sample-Efficient Reinforcement LearningHaoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing LuICLR 2024 · 26 citations
- Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RLQi Lv, Xiang Deng, Gongwei Chen, Michael Yu Wang et al.NeurIPS 2024 · 25 citations
- Rethinking Decision Transformer via Hierarchical Reinforcement LearningYi Ma, Jianye Hao, Hebin Liang, Chenjun XiaoICML 2024 · 15 citations
- Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision TransformersKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2024 · 11 citations
Builds on23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-MakingJincheng Wang, Penny Karanasou, Pengyuan Wei, Elia Gatti et al.ICLR 2025
- Offline Reinforcement Learning with Adaptive Feature FusionTieru Wang, Kunbao Wu, Guoshun NanICLR 2026
- Critic-Guided Decision Transformer for Offline Reinforcement LearningYuanfu Wang, Chao Yang, Ying Wen, Yu Liu et al.AAAI 2024 · 35 citations
- Reinformer: Max-Return Sequence Modeling for Offline RLZifeng Zhuang, Dengyun Peng, Jinxin Liu, Ziqi Zhang et al.ICML 2024 · 29 citations
