Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolution
Qi Tang, Yao Zhao, Meiqin Liu, Jian Jin, Chao Yao
摘要
As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In response to this issue, we introduce a novel paradigm for VSR named Semantic Lens, predicated on semantic priors drawn from degraded videos. Specifically, video is modeled as instances, events, and scenes via a Semantic Extractor. Those semantics assist the Pixel Enhancer in understanding the recovered contents and generating more realistic visual results. The distilled global semantics embody the scene information of each frame, while the instance-specific semantics assemble the spatial-temporal contexts related to each instance. Furthermore, we devise a Semantics-Powered Attention Cross-Embedding (SPACE) block to bridge the pixel-level features with semantic knowledge, composed of a Global Perspective Shifter (GPS) and an Instance-Specific Semantic Embedding Encoder (ISEE). Concretely, the GPS module generates pairs of affine transformation parameters for pixel-level feature modulation conditioned on global semantics. After that the ISEE module harnesses the attention mechanism to align the adjacent frames in the instance-centric semantic space. In addition, we incorporate a simple yet effective pre-alignment module to alleviate the difficulty of model training. Extensive experiments demonstrate the superiority of our model over existing state-of-the-art VSR methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-ResolutionQi Tang, Yao Zhao, Meiqin Liu, Chao YaoNeurIPS 2024 · 被引用 10 次
- Spatial Imputation Drives Cross-Domain Alignment for EEG ClassificationHongjun Liu, Chao Yao, Yalan Zhang, Xiaokun Wang 等ACM MM 2025 · 被引用 1 次
- Event-based Video Super-Resolution via State Space ModelsZeyu Xiao, Xinchao WangCVPR 2025
- Seeing the Unseen: Zooming in the Dark with Event CamerasDachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang 等AAAI 2026
它引用的顶会 Paper16
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan 等NeurIPS 2022 · 被引用 318 次
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee 等NeurIPS 2022 · 被引用 146 次
- Rethinking Alignment in Video Super-Resolution TransformersShuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang 等NeurIPS 2022 · 被引用 134 次
相关 Paper
- SeD: Semantic-Aware Discriminator for Image Super-ResolutionBingchen Li, Xin Li, Hanxin Zhu, Yeying Jin 等CVPR 2024
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-ResolutionHee Min Choi, Hyoa Kang, Suji Kim, Dokwan Oh 等CVPR 2026
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 被引用 113 次
- Event-Enhanced Blurry Video Super-ResolutionDachun Kai, Yueyi Zhang, Jin Wang, Zeyu Xiao 等AAAI 2025 · 被引用 7 次
- EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned LearningYuhan Liu, Linghui Fu, Zhen Yang, Hao Chen 等NeurIPS 2025 · 被引用 3 次
