Enabling True Global Perception in State Space Models for Visual Tasks
Jie Hui, Zhenxiang Zhang, Wenyu Mi, Jianji Wang
摘要
Despite the importance of global contextual modeling in visual tasks, a rigorous mathematical definition remains absent, and the concept is still largely described in heuristic or empirical terms. Existing methods either rely on computationally expensive attention mechanisms or are constrained by the recursive modeling nature of State Space Models (SSMs), making it challenging to achieve both efficiency and true global perception. To address this, we first propose a mathematical definition of global modeling for visual images, providing a theoretical foundation for designing globally-aware and interpretable models. Based on in-depth analysis of SSMs and frequency-domain modeling principles, we construct a complete theoretical framework that overcomes the limitations imposed by SSMs' recursive modeling mechanism from a frequency perspective, thereby adapting SSMs for global perception in image modeling. Guided by this framework, we design the Global-aware SSM (GSSM) module and formally prove that it satisfies definitional requirements of global image modeling. GSSM leverages a Discrete Fourier Transform (DFT)-based modulation mechanism, providing precise front-end control over the SSM's modeling behavior, and enabling efficient global image modeling with linear-logarithmic complexity. Building upon GSSM, we develop GMamba, a plug-and-play module that can be seamlessly integrated at any stage of Convolutional Neural Networks (CNNs). Extensive experiments across multiple tasks, including object detection, semantic segmentation, and instance segmentation, across diverse model architectures, demonstrate that GMamba consistently outperforms existing global modeling modules, validating both the effectiveness of our theoretical framework and the rigor of proposed definition. Code is available at https://github.com/Xinmu-Tantai/GMamba-GSSM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang 等ICLR 2026
- Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language ModelsYujie Feng, Jian Li, Zhihan Zhou, Pengfei Xu 等ICLR 2026
它引用的顶会 Paper32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- AKCMamba-YOLO: Selective State Space Models For Real-Time Object DetectionLong Chen, Hui Wang, Man Xu, Zexuan Li 等CVPR 2026
- GroupMamba: Efficient Group-Based Visual State Space ModelAbdelrahman M. Shaker, Syed Talal Wasim, Salman H. Khan, Juergen Gall 等CVPR 2025
- EfficientVMamba: Atrous Selective Scan for Light Weight Visual MambaXiaohuan Pei, Tao Huang, Chang XuAAAI 2025 · 被引用 248 次
- SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHongda Liu, Longguang Wang, Ye Zhang, Ziru Yu 等CVPR 2025
- Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State FusionChaodong Xiao, Minghan Li, Zhengqiang Zhang, Deyu Meng 等ICLR 2025
