Decision S4: Efficient Sequence-Based RL via State Spaces Layers
Shmuel Bar-David, Itamar Zimerman, Eliya Nachmani, Lior Wolf
摘要
Recently, sequence learning methods have been applied to the problem of off-policy Reinforcement Learning, including the seminal work on Decision Transformers, which employs transformers for this task. Since transformers are parameter-heavy, cannot benefit from history longer than a fixed window size, and are not computed using recurrence, we set out to investigate the suitability of the S4 family of models, which are based on state-space layers and have been shown to outperform transformers, especially in modeling long-range dependencies. In this work we present two main algorithms: (i) an off-policy training procedure that works with trajectories, while still maintaining the training efficiency of the S4 model. (ii) An on-policy training procedure that is trained in a recurrent manner, benefits from long-range dependencies, and is based on a novel stable actor-critic mechanism. Our results indicate that our method outperforms multiple variants of decision transformers, as well as the other baseline methods on most tasks, while reducing the latency, number of parameters, and training time by several orders of magnitude, making our approach more suitable for real-world RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Structured State Space Models for In-Context Reinforcement LearningChris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto 等NeurIPS 2023 · 被引用 164 次
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceHan Zhao, Min Zhang, Wei Zhao, Pengxiang Ding 等AAAI 2025 · 被引用 125 次
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 被引用 53 次
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 被引用 44 次
- Mastering Memory Tasks with World ModelsMohammad Reza Samsami, Artem Zholus, Janarthanan Rajendran, Sarath ChandarICLR 2024 · 被引用 42 次
它引用的顶会 Paper18
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
相关 Paper
- Recurrent Action Transformer with MemoryEgor Cherepanov, Aleksei Staroverov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 被引用 14 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?Yang Dai, Oubo Ma, Longfei Zhang, Xingxing Liang 等NeurIPS 2024 · 被引用 10 次
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li 等NeurIPS 2025
- Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-MakingJincheng Wang, Penny Karanasou, Pengyuan Wei, Elia Gatti 等ICLR 2025
