Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolution
Qi Tang, Yao Zhao, Meiqin Liu, Jian Jin, Chao Yao
Abstract
As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In response to this issue, we introduce a novel paradigm for VSR named Semantic Lens, predicated on semantic priors drawn from degraded videos. Specifically, video is modeled as instances, events, and scenes via a Semantic Extractor. Those semantics assist the Pixel Enhancer in understanding the recovered contents and generating more realistic visual results. The distilled global semantics embody the scene information of each frame, while the instance-specific semantics assemble the spatial-temporal contexts related to each instance. Furthermore, we devise a Semantics-Powered Attention Cross-Embedding (SPACE) block to bridge the pixel-level features with semantic knowledge, composed of a Global Perspective Shifter (GPS) and an Instance-Specific Semantic Embedding Encoder (ISEE). Concretely, the GPS module generates pairs of affine transformation parameters for pixel-level feature modulation conditioned on global semantics. After that the ISEE module harnesses the attention mechanism to align the adjacent frames in the instance-centric semantic space. In addition, we incorporate a simple yet effective pre-alignment module to alleviate the difficulty of model training. Extensive experiments demonstrate the superiority of our model over existing state-of-the-art VSR methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f98d3aa6-27f6-444d-b783-52e1dfd5c303Cited by top-tier papers4
- SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-ResolutionQi Tang, Yao Zhao, Meiqin Liu, Chao YaoNeurIPS 2024 · 10 citations
- Spatial Imputation Drives Cross-Domain Alignment for EEG ClassificationHongjun Liu, Chao Yao, Yalan Zhang, Xiaokun Wang et al.ACM MM 2025 · 1 citation
- Event-based Video Super-Resolution via State Space ModelsZeyu Xiao, Xinchao WangCVPR 2025
- Seeing the Unseen: Zooming in the Dark with Event CamerasDachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang et al.AAAI 2026
Builds on16
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan et al.NeurIPS 2022 · 318 citations
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Rethinking Alignment in Video Super-Resolution TransformersShuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang et al.NeurIPS 2022 · 134 citations
Related papers
- SeD: Semantic-Aware Discriminator for Image Super-ResolutionBingchen Li, Xin Li, Hanxin Zhu, Yeying Jin et al.CVPR 2024
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-ResolutionHee Min Choi, Hyoa Kang, Suji Kim, Dokwan Oh et al.CVPR 2026
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 113 citations
- Event-Enhanced Blurry Video Super-ResolutionDachun Kai, Yueyi Zhang, Jin Wang, Zeyu Xiao et al.AAAI 2025 · 7 citations
- EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned LearningYuhan Liu, Linghui Fu, Zhen Yang, Hao Chen et al.NeurIPS 2025 · 3 citations
