Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling
Sili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang, Tao Luo, Hechang Chen, Lichao Sun, Bo Yang
Abstract
Recent works have shown the remarkable superiority of transformer models in reinforcement learning (RL), where the decision-making problem is formulated as sequential generation. Transformer-based agents could emerge with self-improvement in online environments by providing task contexts, such as multiple trajectories, called in-context RL. However, due to the quadratic computation complexity of attention in transformers, current in-context RL methods suffer from huge computational costs as the task horizon increases. In contrast, the Mamba model is renowned for its efficient ability to process long-term dependencies, which provides an opportunity for in-context RL to solve tasks that require long-term memory. To this end, we first implement Decision Mamba (DM) by replacing the backbone of Decision Transformer (DT). Then, we propose a Decision Mamba-Hybrid (DM-H) with the merits of transformers and Mamba in high-quality prediction and long-term memory. Specifically, DM-H first generates high-value sub-goals from long-term memory through the Mamba model. Then, we use sub-goals to prompt the transformer, establishing high-quality predictions. Experimental results demonstrate that DM-H achieves state-of-the-art in long and short-term tasks, such as D4RL, Grid World, and Tmaze benchmarks. Regarding efficiency, the online testing of DM-H in the long-term task is 28 times faster than the transformer-based baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f76a4f-2a92-4754-92ec-bf50f9c57ba4Cited by top-tier papers10
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?Yang Dai, Oubo Ma, Longfei Zhang, Xingxing Liang et al.NeurIPS 2024 · 10 citations
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 5 citations
- EDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence ModelingJia-Hua Lee, Bor-Jiun Lin, Wei-Fang Sun, Chun-Yi LeeNeurIPS 2025 · 4 citations
- Analytic Energy-Guided Policy Optimization for Offline Reinforcement LearningJifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu et al.NeurIPS 2025 · 4 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
Related papers
- In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-ThoughtSili Huang, Jifeng Hu, Hechang Chen, Lichao Sun et al.ICML 2024 · 25 citations
- QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RLXing Lei, Jincheng Wang, Xuetao Zhang, Donglin WangICML 2026
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal et al.NeurIPS 2025 · 9 citations
- Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter EfficientWenlong Wang, Ivana Dusparic, Yucheng Shi, Ke Zhang et al.ICLR 2025
- DeciMamba: Exploring the Length Extrapolation Potential of MambaAssaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen et al.ICLR 2025
