Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-Making
Jincheng Wang, Penny Karanasou, Pengyuan Wei, Elia Gatti, Diego Martínez Plasencia, Dimitrios Kanoulas
Abstract
Decision Transformers (DTs) effectively capture long-range dependencies using self-attention but struggle with fine-grained local relationships, especially the Markovian properties in many offline-RL datasets. Conversely, Decision ConvFormer (DC) utilizes convolutional filters for capturing local patterns but shows limitations in tasks demanding long-term dependencies, such as Maze2d. To address these limitations and leverage both strengths, we propose the Long-Short Decision Transformer (LSDT), a general-purpose architecture to effectively capture global and local dependencies across two specialized parallel branches (self-attention and convolution). We explore how these branches complement each other by modeling various ranged dependencies across different environments, and compare it against other baselines. Experimental results demonstrate our LSDT achieves state-of-the-art performance and notable gains over the standard DT in D4RL offline RL benchmark. Leveraging the parallel architecture, LSDT performs consistently on diverse datasets, including Markovian and non-Markovian. We also demonstrate the flexibility of LSDT's architecture, where its specialized branches can be replaced or integrated into models like DC to improve performance in capturing diverse dependencies. Finally, we also highlight the role of goal states in improving decision-making for goal-reaching tasks like Antmaze. We conduct evaluations of our proposed approaches on the standard D4RL benchmark (Fu et al., 2020) . Results from our studies demonstrate consistent enhancement and superior performance of our LSDT in decision-making compared to DT and its variants. LSDT achieves comparable performance compared to state-of-the-art RL methods. Additionally, the goal-state conditioning method significantly enhances the performance of transformer-based RL, as shown in Antmaze and Maze2d tasks. To validate the flexibility of LSDT, we replaced the Dynamic Conv in the short-term branch with DC's filters, and the observed improvements (e.g., Figure 2 ) on non-Markovian datasets confirm that our approach is straightforward and plug-and-play with minimal adjustments. PRELIMINARIES Offline Reinforcement Learning. In RL, the learning environment can be considered as a Markov Decision Process (MDP), defined by tuples (S, A, P, R). Here, S represents the set of possible states, A represents the set of actions, P represents the state transition probabilities P (s ′ |s, a), and R
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Recurrent Action Transformer with MemoryEgor Cherepanov, Aleksei Staroverov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 14 citations
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent FrameworkShayan Talaei, Meijin Li, Kanu Grover, James Kent Hippler et al.UIST 2025 · 4 citations
- Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised LearningYuxiao Yang, Weitong ZhangICML 2026
- Decision Mixer: Integrating Long-term and Local Dependencies via Dynamic Token Selection for Decision-MakingHongling Zheng, Li Shen, Yong Luo, Deheng Ye et al.ICML 2025
- QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RLXing Lei, Jincheng Wang, Xuetao Zhang, Donglin WangICML 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
Related papers
- Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision MakingJeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul SungICLR 2024 · 36 citations
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li et al.NeurIPS 2025
- Unveiling Markov heads in Pretrained Language Models for Offline Reinforcement LearningWenhao Zhao, Qiushui Xu, Linjie Xu, Lei Song et al.ICML 2025
- Offline Reinforcement Learning with Adaptive Feature FusionTieru Wang, Kunbao Wu, Guoshun NanICLR 2026
