VSSD: Vision Mamba With Non-Causal State Space Duality
Yuheng Shi, Mingjia Li, Minjing Dong, Chang Xu
Abstract
Vision transformers have significantly advanced the field of computer vision, offering robust modeling capabilities and global receptive field. However, their high computational demands limit their applicability in processing long sequences. To tackle this issue, State Space Models (SSMs) have gained prominence in vision tasks as they offer linear computational complexity. Recently, State Space Duality (SSD), an improved variant of SSMs, was introduced in Mamba2 to enhance model performance and efficiency. However, the inherent causal nature of SSD/SSMs restricts their applications in non-causal vision tasks. To address this limitation, we introduce Visual State Space Duality (VSSD) model, which has a non-causal format of SSD. Specifically, we propose to discard the magnitude of interactions between the hidden state and tokens while preserving their relative weights, which relieves the dependencies of token contribution on previous tokens. Together with the involvement of multi-scan strategies, we show that the scanning results can be integrated to achieve non-causality, which not only improves the performance of SSD in vision tasks but also enhances its efficiency. We conduct extensive experiments on various benchmarks including image classification, detection, and segmentation, where VSSD surpasses existing state-of-the-art SSM-based models. Code and weights are available at https://github.com/YuHengsss/VSSD.
Recently, State Space Models (SSMs) [15,17,16,49], exemplified by Mamba [14], have garnered considerable attention from researchers. The S6 block, in particular, offers a global receptive field and exhibits linear complexity with respect to sequence length, presenting an efficient alternative. Pioneering vision mamba models such as Vim [69] and VMamba [35] have been developed to apply SSMs to vision tasks. Afterward, many variants were proposed [29,50,61,47], which flatten 2D feature maps into 1D sequences using different scanning routes, model them with the S6 block, and subsequently integrate the results in multiple scanning routes. These multi-scan approaches improve Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd7a2e4e-96f9-43b8-97d0-fe65dcc7483fCited by top-tier papers7
- ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial ImagesXilei Zhu, Liu Yang, Huiyu Duan, Xiongkuo Min et al.IEEE VR 2025 · 10 citations
- MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image RecoveryHainuo Wang, Qiming Hu, Xiaojie GuoNeurIPS 2025 · 8 citations
- PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationFei Xie, Zhongdao Wang, Weijia Zhang, Chao MaICCV 2025 · 2 citations
- Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation ModelsChristian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito et al.CVPR 2026 · 2 citations
- Partial Ring Scan: Revisiting Scan Order in Vision State Space ModelsYi-Kuan Hsieh, Kuan-Chuan Peng, Xin Li, Ming-Ching Chang et al.ICML 2026
Builds on38
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space ModelYuheng Shi, Minjing Dong, Chang XuNeurIPS 2024 · 129 citations
- EfficientVMamba: Atrous Selective Scan for Light Weight Visual MambaXiaohuan Pei, Tao Huang, Chang XuAAAI 2025 · 248 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Boosting Vision State Space Model with Fractal ScanningHaoke Xiao, Lv Tang, Peng-Tao Jiang, Hao Zhang et al.AAAI 2025 · 8 citations
- EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space DualitySanghyeok Lee, Joonmyung Choi, Hyunwoo J. KimCVPR 2025
