Event-based Video Super-Resolution via State Space Models
Zeyu Xiao, Xinchao Wang
Abstract
Exploiting temporal correlations is crucial for video super-resolution (VSR). Recent approaches enhance this by incorporating event cameras. In this paper, we introduce MamEVSR, a Mamba-based network for event-based VSR that leverages the selective state space model, Mamba. MamEVSR stands out by offering global receptive field coverage with linear computational complexity, thus addressing the limitations of convolutional neural networks and Transformers. The key components of MamEVSR include: (1) The interleaved Mamba (iMamba) block, which interleaves tokens from adjacent frames and applies multidirectional selective state space modeling, enabling efficient feature fusion and propagation across bi-directional frames while maintaining linear complexity. (2) The crossmodality Mamba (cMamba) block facilitates further interaction and aggregation between event information and the output from the iMamba block. The cMamba block can leverage complementary spatio-temporal information from both modalities and allows MamEVSR to capture finer motion details. Experimental results show that the proposed MamEVSR achieves superior performance on various datasets quantitatively and qualitatively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ba334f4-d5dc-4c48-be25-c425c33f5377Cited by top-tier papers10
- A Unified Solution to Video Fusion: From Multi-Frame Learning to BenchmarkingZixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui et al.NeurIPS 2025 · 21 citations
- ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image RestorationXu Zhang, Huan Zhang, Guoli Wang, Qian Zhang et al.AAAI 2026 · 6 citations
- Event6D: Event-based Novel Object 6D Pose TrackingJae-Young Kang, Hoonhee Cho, Taeyeop Lee, Minjun Kang et al.CVPR 2026 · 4 citations
- Trajectory-aware Shifted State Space Models for Online Video Super-ResolutionQiang Zhu, Xiandong Meng, Yuxuan Jiang, Fan Zhang et al.ICLR 2026 · 3 citations
- Seeing the Unseen: Zooming in the Dark with Event CamerasDachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang et al.AAAI 2026
Builds on35
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- EVDM: Event-based Real-World Video Deblurring with MambaZhijing Sun, Senyan Xu, Kean Liu, Runze Tian et al.ICCV 2025 · 6 citations
- EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video ReconstructionChengjie Ge, Xueyang Fu, Peng He, Kunyu Wang et al.AAAI 2025 · 6 citations
- VSRM: A Robust Mamba-Based Framework for Video Super-ResolutionDinh Phu Tran, Dao Duy Hung, Daeyoung KimICCV 2025 · 4 citations
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationRunyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim et al.ICCV 2025 · 2 citations
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 15 citations
