Spectral State Space Model for Rotation-Invariant Visual Representation Learning
Sahar Dastani, Ali Bahri, Moslem Yazdanpanah, Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Farzad Beizaee, Milad Cheraghalikhani, Arnab Kumar Mondal, Herve Lombaert, Christian Desrosiers
Abstract
State Space Models (SSMs) have recently emerged as an alternative to Vision Transformers (ViTs) due to their unique ability of modeling global relationships with linear complexity. SSMs are specifically designed to capture spatially proximate relationships of image patches. However, they fail to identify relationships between conceptually related yet not adjacent patches. This limitation arises from the non-causal nature of image data, which lacks inherent directional relationships. Additionally, current vision-based SSMs are highly sensitive to transformations such as rotation. Their predefined scanning directions depend on the original image orientation, which can cause the model to produce inconsistent patch-processing sequences after rotation. To address these limitations, we introduce Spectral VMamba, a novel approach that effectively captures the global structure within an image by leveraging spectral information derived from the graph Laplacian of image patches. Through spectral decomposition, our approach encodes patch relationships independently of image orientation, achieving patch traversal rotation invariance with our Rotational Feature Normalizer (RFN) module. Our experiments on classification tasks show that Spectral VMamba outperforms the leading SSM models in vision, such as VMamba, while maintaining invariance to rotations and a providing a similar runtime efficiency. The implementation is available at: https://github.com/ Sahardastani/spectral_vmamba.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59c81d4b-6e2b-4f58-ba3f-f9f58badc83dCited by top-tier papers3
- Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token PurgingMoslem Yazdanpanah, Ali Bahri, Mehrdad Noori, Sahar Dastani et al.ICCV 2025 · 3 citations
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM TraversesSahar Dastani, Ali Bahri, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2025
- 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information PropagationLonglong Yu, Wenxi Li, Yaoqi Sun, Hang Xu et al.AAAI 2026
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
Related papers
- VSSD: Vision Mamba With Non-Causal State Space DualityYuheng Shi, Mingjia Li, Minjing Dong, Chang XuICCV 2025 · 20 citations
- EfficientVMamba: Atrous Selective Scan for Light Weight Visual MambaXiaohuan Pei, Tao Huang, Chang XuAAAI 2025 · 248 citations
- DAMamba: Vision State Space Model with Dynamic Adaptive ScanTanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei et al.NeurIPS 2025 · 24 citations
- Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State FusionChaodong Xiao, Minghan Li, Zhengqiang Zhang, Deyu Meng et al.ICLR 2025
- Partial Ring Scan: Revisiting Scan Order in Vision State Space ModelsYi-Kuan Hsieh, Kuan-Chuan Peng, Xin Li, Ming-Ching Chang et al.ICML 2026
