Efficient Visual State Space Model for Image Deblurring
Lingshun Kong, Jiangxin Dong, Jinhui Tang, Ming-Hsuan Yang, Jinshan Pan
Abstract
Convolutional neural networks (CNNs) and Vision Transformers (ViTs) have achieved excellent performance in image restoration. While ViTs generally outperform CNNs by effectively capturing long-range dependencies and inputspecific characteristics, their computational complexity increases quadratically with image resolution. This limitation hampers their practical application in high-resolution image restoration. In this paper, we propose a simple yet effective visual state space model (EVSSM) for image deblurring, leveraging the benefits of state space models (SSMs) for visual data. In contrast to existing methods that employ several fixed-direction scanning for feature extraction, which significantly increases the computational cost, we develop an efficient visual scan block that applies various geometric transformations before each SSM-based module, capturing useful non-local information and maintaining high efficiency. In addition, to more effectively capture and represent local information, we propose an efficient discriminative frequency domain-based feedforward network (EDFFN), which can effectively estimate useful frequency information for latent clear image restoration. Extensive experimental results show that the proposed EVSSM performs favorably against state-of-the-art methods on benchmark datasets and real-world images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d7dd0ae-bd3b-497b-b20f-67aeba3e6de5Cited by top-tier papers11
- 4KAgent: Agentic Any Image to 4K Super-ResolutionYushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang et al.NeurIPS 2025 · 51 citations
- LoFormer: Local Frequency Transformer for Image DeblurringXintian Mao, Jiansheng Wang, Xingran Xie, Qingli Li et al.ACM MM 2024 · 44 citations
- Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image RestorationChen Wu, Ling Wang, Zhuoran Zheng, Yuning Cui et al.CVPR 2026 · 9 citations
- MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image RecoveryHainuo Wang, Qiming Hu, Xiaojie GuoNeurIPS 2025 · 8 citations
- EVDM: Event-based Real-World Video Deblurring with MambaZhijing Sun, Senyan Xu, Kean Liu, Runze Tian et al.ICCV 2025 · 6 citations
Builds on23
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- FFA-Net: Feature Fusion Attention Network for Single Image DehazingXu Qin, Zhilin Wang, Yuanchao Bai, Xiaodong Xie et al.AAAI 2020 · 1,828 citations
Related papers
- Learning Enriched Features via Selective State Spaces Model for Efficient Image DeblurringHu Gao, Bowen Ma, Ying Zhang, Jingfan Yang et al.ACM MM 2024 · 25 citations
- VSSD: Vision Mamba With Non-Causal State Space DualityYuheng Shi, Mingjia Li, Minjing Dong, Chang XuICCV 2025 · 20 citations
- DAMamba: Vision State Space Model with Dynamic Adaptive ScanTanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei et al.NeurIPS 2025 · 24 citations
- Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space ModelYuheng Shi, Minjing Dong, Chang XuNeurIPS 2024 · 129 citations
- Efficient Frequency Domain-based Transformers for High-Quality Image DeblurringLingshun Kong, Jiangxin Dong, Jianjun Ge, Mingqiang Li et al.CVPR 2023
