When should we prefer Decision Transformers for Offline Reinforcement Learning?
Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani, Amy Zhang
Abstract
Offline reinforcement learning (RL) allows agents to learn effective, return-maximizing policies from a static dataset. Three popular algorithms for offline RL are Conservative Q-Learning (CQL), Behavior Cloning (BC), and Decision Transformer (DT), from the class of Q-Learning, Imitation Learning, and Sequence Modeling respectively. A key open question is: which algorithm is preferred under what conditions? We study this question empirically by exploring the performance of these algorithms across the commonly used D4RL and Robomimic benchmarks. We design targeted experiments to understand their behavior concerning data suboptimality, task complexity, and stochasticity. Our key findings are: (1) DT requires more data than CQL to learn competitive policies but is more robust; (2) DT is a substantially better choice than both CQL and BC in sparse-reward and low-quality data settings; (3) DT and BC are preferable as task horizon increases, or when data is obtained from human demonstrators; and (4) CQL excels in situations characterized by the combination of high stochasticity and low data quality. We also investigate architectural choices and scaling trends for DT on Atari and D4RL and make design/scaling recommendations. We find that scaling the amount of data for DT by 5x gives a 2.5x average score improvement on Atari.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu et al.NeurIPS 2024 · 31 citations
- Multi-agent Coordination via Flow MatchingDongsu Lee, Daehee Lee, Amy ZhangICLR 2026 · 9 citations
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu et al.NeurIPS 2025 · 1 citation
- Structured Expert Routing with Multi-View Task Priors for Offline Meta-Reinforcement LearningYisen Zhao, Peixi Peng, Xinyu Hu, Cong Li et al.ICML 2026
- Tackling Data Corruption in Offline Reinforcement Learning via Sequence ModelingJiawei Xu, Rui Yang, Shuang Qiu, Feng Luo et al.ICLR 2025
Builds on21
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
- Rethinking Decision Transformer via Hierarchical Reinforcement LearningYi Ma, Jianye Hao, Hebin Liang, Chenjun XiaoICML 2024 · 15 citations
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar et al.NeurIPS 2023 · 34 citations
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 99 citations
