Rethinking Token Reduction for State Space Models
Zheng Zhan, Yushu Wu, Zhenglun Kong, Changdi Yang, Yifan Gong, Xuan Shen, Xue Lin, Pu Zhao, Yanzhi Wang
Abstract
Recent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies. Architectures like Mamba have scaled to billions of parameters with selective SSM. To facilitate broader applications using Mamba, exploring its efficiency is crucial. While token reduction techniques offer a straightforward post-training strategy, we find that applying existing methods directly to SSMs leads to substantial performance drops. Through insightful analysis, we identify the reasons for this failure and the limitations of current techniques. In response, we propose a tailored, unified post-training token reduction method for SSMs. Our approach integrates token importance and similarity, thus taking advantage of both pruning and merging, to devise a fine-grained intra-layer token reduction strategy. Extensive experiments show that our method improves the average accuracy by 5.7% to 13.1% on six benchmarks with Mamba-2 compared to existing methods, while significantly reducing computational demands and memory requirements. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Efficient Reasoning with Hidden ThinkingXuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi et al.ICML 2026 · 56 citations
- LazyDiT: Lazy Learning for the Acceleration of Diffusion TransformersXuan Shen, Zhao Song, Yufa Zhou, Bo Chen et al.AAAI 2025 · 40 citations
- Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced CoveringYangfu Li, Hongjian Zhan, Tianyi Chen, Qi Liu et al.NeurIPS 2025 · 11 citations
- Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning PathsJindi Lv, Yuhao Zhou, Mingjia Shi, Zhiyuan Liang et al.NeurIPS 2025
- Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space ModelsJindi Lv, Aoyu Li, Yuhao Zhou, Zheng Zhu et al.ICML 2026
Builds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- Exploring Token Pruning in Vision State Space ModelsZheng Zhan, Zhenglun Kong, Yifan Gong, Yushu Wu et al.NeurIPS 2024 · 34 citations
- SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-IdentificationHuiyuan Huang, SANG MIN YOONCVPR 2026
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained EnvironmentsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaEMNLP 2025 · 1 citation
- Demystifying the Token Dynamics of Deep Selective State Space ModelsThieu N. Vo, Duy-Tung Pham, Xin T. Tong, Tan Minh NguyenICLR 2025
- LongMamba: Enhancing Mamba's Long-Context Capabilities via Training-Free Receptive Field EnlargementZhifan Ye, Kejing Xia, Yonggan Fu, Xin Dong et al.ICLR 2025
