Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
Jindi Lv, Yuhao Zhou, Mingjia Shi, Zhiyuan Liang, Panpan Zhang, Xiaojiang Peng, Wangbo Zhao, Zheng Zhu, Jiancheng Lv, Qing Ye, Kai Wang
Abstract
Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields. In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness. Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages. and merging [18,19]) for Mamba have garnered increasing attention as promising avenues toward further optimization.
Nevertheless, Mamba exhibits significantly greater performance degradation during token reduction compared to Transformers (eg., ViT [13]), as shown in Figure 1a. This discrepancy arises from Transformers using self-attention to establish fully connected token relationships, while Mamba processes tokens sequentially along chain-like scanning paths. This chain-based structure makes Mamba susceptible to cascading information loss during token reduction. For clarity, an illustrative example is provided in Figure 1b.
Recent Mamba variants [20,21,22], such as Vim [6], have attempted to mitigate this limitation through dual-path scanning strategies that combine forward and reverse sequential paths. While these symmetric designs enhance sequence modeling capabilities and improve baseline performance, they remain ineffective for token reduction. The inherent symmetry of dual-path scanning confines token relationships to the same chain-like structure, failing to address the systemic vulnerability to large-scale connection disruption during pruning.
Inspired by these observations, we hypothesize that minimizing connection disruption during token reduction can mitigate performance drop. To validate this, we introduce asymmetric scanning into dual-path Mamba. As illustrated in Figure 1c, asymmetric scanning paths reduce accuracy degradation from 38% (with symmetric paths) to 21% at the same pruning ratio. This suggests that diversifying chain-like dependencies effectively mitigate pruning-induced performance decline.
To further elucidate this phenomenon, we quantify the token connection survival rate across different dual-path strategies during token reduction. Figure 1c reveals a strong positive correlation between connection survival rates and accuracy, with asymmetric paths exhibiting superior robustness. This confirms our hypothesis: enhancing token connection survival rates via asymmetric path diversification is pivotal for improving Mamba's pruning resilience.
In this work, we propose AMVim, a novel Asymmetric Multi-scale Vision Mamba for pruning robustness. To enhance space information diversity, we integrate a window-aware scanning mechanism into one path. By adopting a different scanning direction within windows compared to the main path, we construct multi-level asymmetric paths. This multi-dimensional information flow enables each token to perceive neighborhood information from multiple perspectives. Furthermore, the integration of window-based scanning with the main path creates a multi-level complementary design, allowing for interactions between global context and local dependencies.
Empirically, as shown in Figure 1c, our method results in just a 3% accuracy drop during token reduction (with the blue dotted line representing the baseline accuracy of 76.1%) on ImageNet-1K, achieving a 34% improvement over Vim. This highlights that the multi-scale scanning mechanism enhances the spatial awareness of SSMs, significantly reducing token sensitivity to local variations.
We highlight the main contributions of this paper below:
• We hypothesize token connection survival rate is a critical factor in performance degradation and propose asymmetric scanning paths to effectively mitigate this issue.
• We design a multi-scale asymmetric scanning mechanism that balances global and local spatial information while preserving the benefits of asymmetric paths.
• Our method achieves state-of-the-art pruning resilience, outperforming Vim-T by 34% on ImageNet-1K under identical parameters and FLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22d3c1bf-a90c-4d5b-ace8-c672b62a50feCited by top-tier papers1
Ask how each one uses itBuilds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
Related papers
- SF-Mamba: Rethinking State Space Model for VisionMasakazu Yoshimura, Teruaki Hayashi, Yuki Hoshino, Wei-Yao Wang et al.ICML 2026
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei et al.CVPR 2025
- Exploring Token Pruning in Vision State Space ModelsZheng Zhan, Zhenglun Kong, Yifan Gong, Yushu Wu et al.NeurIPS 2024 · 34 citations
- QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space ModelFei Xie, Weijia Zhang, Zhongdao Wang, Chao MaNeurIPS 2024 · 40 citations
- Stochastic Layer-Wise Shuffle for Improving Vision Mamba TrainingZizheng Huang, Haoxing Chen, Jiaqi Li, Jun Lan et al.ICML 2025
