Lune

ICCV2025Top-tier venue

VSSD: Vision Mamba With Non-Causal State Space Duality

Yuheng Shi, Mingjia Li, Minjing Dong, Chang Xu

2025Year
20Citations
7Top-tier citations

Abstract

Vision transformers have significantly advanced the field of computer vision, offering robust modeling capabilities and global receptive field. However, their high computational demands limit their applicability in processing long sequences. To tackle this issue, State Space Models (SSMs) have gained prominence in vision tasks as they offer linear computational complexity. Recently, State Space Duality (SSD), an improved variant of SSMs, was introduced in Mamba2 to enhance model performance and efficiency. However, the inherent causal nature of SSD/SSMs restricts their applications in non-causal vision tasks. To address this limitation, we introduce Visual State Space Duality (VSSD) model, which has a non-causal format of SSD. Specifically, we propose to discard the magnitude of interactions between the hidden state and tokens while preserving their relative weights, which relieves the dependencies of token contribution on previous tokens. Together with the involvement of multi-scan strategies, we show that the scanning results can be integrated to achieve non-causality, which not only improves the performance of SSD in vision tasks but also enhances its efficiency. We conduct extensive experiments on various benchmarks including image classification, detection, and segmentation, where VSSD surpasses existing state-of-the-art SSM-based models. Code and weights are available at https://github.com/YuHengsss/VSSD.

Recently, State Space Models (SSMs) [15,17,16,49], exemplified by Mamba [14], have garnered considerable attention from researchers. The S6 block, in particular, offers a global receptive field and exhibits linear complexity with respect to sequence length, presenting an efficient alternative. Pioneering vision mamba models such as Vim [69] and VMamba [35] have been developed to apply SSMs to vision tasks. Afterward, many variants were proposed [29,50,61,47], which flatten 2D feature maps into 1D sequences using different scanning routes, model them with the S6 block, and subsequently integrate the results in multiple scanning routes. These multi-scan approaches improve Preprint. Under review.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cd7a2e4e-96f9-43b8-97d0-fe65dcc7483f

Cited by top-tier papers7

Ask how each one uses it

Builds on38

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines