DeciMamba: Exploring the Length Extrapolation Potential of Mamba
Assaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen, Amir Globerson, Lior Wolf, Raja Giryes
摘要
Long-range sequence processing poses a significant challenge for Transformers due to their quadratic complexity in input length. A promising alternative is Mamba, which demonstrates high performance and achieves Transformer-level capabilities while requiring substantially fewer computational resources. In this paper we explore the length-generalization capabilities of Mamba, which we find to be relatively limited. Through a series of visualizations and analyses we identify that the limitations arise from a restricted effective receptive field, dictated by the sequence length used during training. To address this constraint, we introduce DeciMamba, a context-extension method specifically designed for Mamba. This mechanism, built on top of a hidden filtering mechanism embedded within the S6 layer, enables the trained model to extrapolate well even without additional training. Empirical experiments over real-world long-range NLP tasks show that DeciMamba can extrapolate to context lengths that are significantly longer than the ones seen during training, while enjoying faster inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes 等NeurIPS 2024 · 被引用 14 次
- MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba AttentionZilong Zhao, Zhengming Ding, Pei Niu, Wenhao Sun 等CVPR 2026 · 被引用 12 次
- Improving Bilinear RNN with Closed-loop ControlJiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan 等NeurIPS 2025 · 被引用 10 次
- Long-Context State-Space Video World ModelsRyan Po, Yotam Nitzan, Richard Zhang, Berlin Chen 等ICCV 2025 · 被引用 6 次
- Hymba: A Hybrid-head Architecture for Small Language ModelsXin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon 等ICLR 2025 · 被引用 2 次
它引用的顶会 Paper32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- MambaExtend: A Training-Free Approach to Improve Long Context Extension of MambaSeyedarmin Azizi, Souvik Kundu, Mohammad Erfan Sadeghi, Massoud PedramICLR 2025
- LongMamba: Enhancing Mamba's Long-Context Capabilities via Training-Free Receptive Field EnlargementZhifan Ye, Kejing Xia, Yonggan Fu, Xin Dong 等ICLR 2025
- Mamba Modulation: On the Length Generalization of Mamba ModelsPeng Lu, Jerry Huang, Qiuhao Zeng, Xinyu Wang 等NeurIPS 2025 · 被引用 2 次
- Exploring the Limitations of Mamba in COPY and CoT ReasoningRuifeng Ren, Zhicong Li, Yong LiuEMNLP 2025 · 被引用 6 次
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky 等ICML 2026
