Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
Xingyu Zhou, Leheng Zhang, Xiaorui Zhao, Keze Wang, Leida Li, Shuhang Gu
Abstract
Recently, Vision Transformer has achieved great success in recovering missing details in low-resolution sequences, i.e., the video super-resolution (VSR) task. Despite its su-periority in VSR accuracy, the heavy computational bur-den as well as the large memory footprint hinder the de-ployment of Transformer-based VSR models on constrained devices. In this paper, we address the above issue by proposing a novel feature-level masked processing frame-work: VSR with Masked Intra and inter-frame Attention (MIA-VSR). The core of MIA-VSR is leveraging feature-level temporal continuity between adjacent frames to re-duce redundant computations and make more rational use of previously enhanced SR features. Concretely, we propose an intra-frame and inter-frame attention block which takes the respective roles of past features and input features into consideration and only exploits previously enhanced fea-tures to provide supplementary information. In addition, an adaptive block-wise mask prediction module is developed to skip unimportant computations according to feature sim-ilarity between adjacent frames. We conduct detailed ab-lation studies to validate our contributions and compare the proposed method with recent state-of-the-art VSR approaches. The experimental results demonstrate that MIA-VSR improves the memory and computation efficiency over state-of-the-art methods, without trading off PSNR accuracy. The code is available at https://github.com/LabShuHangGU/MIA-VSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcde4e59-1899-40f0-b3a4-769348f296b1Cited by top-tier papers21
- PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementZhanfeng Feng, Long Peng, Xin Di, Yong Guo et al.NeurIPS 2025 · 17 citations
- Event-Enhanced Blurry Video Super-ResolutionDachun Kai, Yueyi Zhang, Jin Wang, Zeyu Xiao et al.AAAI 2025 · 7 citations
- VSRM: A Robust Mamba-Based Framework for Video Super-ResolutionDinh Phu Tran, Dao Duy Hung, Daeyoung KimICCV 2025 · 4 citations
- HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionYang Zou, Xingyue Zhu, Kaiqi Han, Jun Ma et al.AAAI 2026 · 3 citations
- Trajectory-aware Shifted State Space Models for Online Video Super-ResolutionQiang Zhu, Xiandong Meng, Yuxuan Jiang, Fan Zhang et al.ICLR 2026 · 3 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan et al.NeurIPS 2022 · 318 citations
- Self-Guided Network for Fast Image DenoisingShuhang Gu, Yawei Li, Luc Van Gool, Radu TimofteICCV 2019 · 187 citations
- Rethinking Alignment in Video Super-Resolution TransformersShuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang et al.NeurIPS 2022 · 134 citations
Related papers
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 113 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- How Video Super-Resolution and Frame Interpolation Mutually BenefitChengcheng Zhou, Zongqing Lu, Linge Li, Qiangyu Yan et al.ACM MM 2021 · 12 citations
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision TransformersBowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang et al.NeurIPS 2021 · 209 citations
- InterArch: Video Transformer Acceleration via Inter-Feature Deduplication with Cube-based DataflowXuhang Wang, Zhuoran Song, Xiaoyao LiangDAC 2024 · 2 citations
