Rethinking Decision Transformer via Hierarchical Reinforcement Learning
Yi Ma, Jianye Hao, Hebin Liang, Chenjun Xiao
Abstract
Decision Transformer (DT) is an innovative algorithm leveraging recent advances of the transformer architecture in reinforcement learning (RL). However, a notable limitation of DT is its reliance on recalling trajectories from datasets, losing the capability to seamlessly stitch sub-optimal trajectories together. In this work we introduce a general sequence modeling framework for studying sequential decision making through the lens of Hierarchical RL. At the time of making decisions, a high-level policy first proposes an ideal prompt for the current state, a low-level policy subsequently generates an action conditioned on the given prompt. We show DT emerges as a special case of this framework with certain choices of high-level and low-level policies, and discuss the potential failure of these choices. Inspired by these observations, we study how to jointly optimize the high-level and low-level policies to enable the stitching ability, which further leads to the development of new offline RL algorithms. Our empirical results clearly show that the proposed algorithms significantly surpass DT on several control and navigation benchmarks. We hope our contributions can inspire the integration of transformer architectures within the field of RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a331d9f2-38fd-4879-8d25-f6a225e5cfb9Cited by top-tier papers8
- Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RLQi Lv, Xiang Deng, Gongwei Chen, Michael Yu Wang et al.NeurIPS 2024 · 25 citations
- Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision TransformersKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2024 · 11 citations
- PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-PerformerChang Chen, Junyeob Baek, Fei Deng, Kenji Kawaguchi et al.ICML 2024 · 4 citations
- Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement LearningYixiu Mao, Yun Qu, Qi (Cheems) Wang, Xiangyang JiNeurIPS 2025 · 3 citations
- Ad Hoc Teamwork via Offline Goal-Based Decision TransformersXinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai et al.ICML 2025
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
Related papers
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Reinformer: Max-Return Sequence Modeling for Offline RLZifeng Zhuang, Dengyun Peng, Jinxin Liu, Ziqi Zhang et al.ICML 2024 · 29 citations
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li et al.NeurIPS 2025
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu et al.ICML 2022 · 194 citations
