VSSD: Vision Mamba With Non-Causal State Space Duality
Yuheng Shi, Mingjia Li, Minjing Dong, Chang Xu
摘要
Vision transformers have significantly advanced the field of computer vision, offering robust modeling capabilities and global receptive field. However, their high computational demands limit their applicability in processing long sequences. To tackle this issue, State Space Models (SSMs) have gained prominence in vision tasks as they offer linear computational complexity. Recently, State Space Duality (SSD), an improved variant of SSMs, was introduced in Mamba2 to enhance model performance and efficiency. However, the inherent causal nature of SSD/SSMs restricts their applications in non-causal vision tasks. To address this limitation, we introduce Visual State Space Duality (VSSD) model, which has a non-causal format of SSD. Specifically, we propose to discard the magnitude of interactions between the hidden state and tokens while preserving their relative weights, which relieves the dependencies of token contribution on previous tokens. Together with the involvement of multi-scan strategies, we show that the scanning results can be integrated to achieve non-causality, which not only improves the performance of SSD in vision tasks but also enhances its efficiency. We conduct extensive experiments on various benchmarks including image classification, detection, and segmentation, where VSSD surpasses existing state-of-the-art SSM-based models. Code and weights are available at https://github.com/YuHengsss/VSSD.
Recently, State Space Models (SSMs) [15,17,16,49], exemplified by Mamba [14], have garnered considerable attention from researchers. The S6 block, in particular, offers a global receptive field and exhibits linear complexity with respect to sequence length, presenting an efficient alternative. Pioneering vision mamba models such as Vim [69] and VMamba [35] have been developed to apply SSMs to vision tasks. Afterward, many variants were proposed [29,50,61,47], which flatten 2D feature maps into 1D sequences using different scanning routes, model them with the S6 block, and subsequently integrate the results in multiple scanning routes. These multi-scan approaches improve Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial ImagesXilei Zhu, Liu Yang, Huiyu Duan, Xiongkuo Min 等IEEE VR 2025 · 被引用 10 次
- MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image RecoveryHainuo Wang, Qiming Hu, Xiaojie GuoNeurIPS 2025 · 被引用 8 次
- PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationFei Xie, Zhongdao Wang, Weijia Zhang, Chao MaICCV 2025 · 被引用 2 次
- Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation ModelsChristian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito 等CVPR 2026 · 被引用 2 次
- Partial Ring Scan: Revisiting Scan Order in Vision State Space ModelsYi-Kuan Hsieh, Kuan-Chuan Peng, Xin Li, Ming-Ching Chang 等ICML 2026
它引用的顶会 Paper38
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space ModelYuheng Shi, Minjing Dong, Chang XuNeurIPS 2024 · 被引用 129 次
- EfficientVMamba: Atrous Selective Scan for Light Weight Visual MambaXiaohuan Pei, Tao Huang, Chang XuAAAI 2025 · 被引用 248 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Boosting Vision State Space Model with Fractal ScanningHaoke Xiao, Lv Tang, Peng-Tao Jiang, Hao Zhang 等AAAI 2025 · 被引用 8 次
- EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space DualitySanghyeok Lee, Joonmyung Choi, Hyunwoo J. KimCVPR 2025
