Learning Trajectory-Aware Transformer for Video Super-Resolution
Chengxu Liu, Huan Yang, Jianlong Fu, Xueming Qian
摘要
Video super-resolution (VSR) aims to restore a sequence of high-resolution (HR) frames from their low-resolution (LR) counterparts. Although some progress has been made, there are grand challenges to effectively utilize temporal dependency in entire video sequences. Existing approaches usually align and aggregate video frames from limited adjacent frames (e.g., 5 or 7 frames), which prevents these approaches from satisfactory results. In this paper, we take one step further to enable effective spatio-temporal learning in videos. We propose a novel Trajectory-aware Transformer for Video Super-Resolution (TTVSR). In particular, we formulate video frames into several pre-aligned trajectories which consist of continuous visual tokens. For a query token, self-attention is only learned on relevant visual tokens along spatio-temporal trajectories. Compared with vanilla vision Transformers, such a design significantly reduces the computational cost and enables Transformers to model long-range features. We further propose a cross-scale feature tokenization module to over-come scale-changing problems that often occur in long-range videos. Experimental results demonstrate the superiority of the proposed TTVSR over state-of-the-art models, by extensive quantitative and qualitative evaluations in four widely-used video super-resolution benchmarks. Both code and pre-trained models can be downloaded at https://github.com/researchmm/TTVSR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan 等NeurIPS 2022 · 被引用 318 次
- Zero-Reference Low-Light Enhancement via Physical Quadruple PriorsWenjing Wang, Huan Yang, Jianlong Fu, Jiaying LiuCVPR 2024 · 被引用 90 次
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li 等CVPR 2026 · 被引用 42 次
- Learning Truncated Causal History Model for Video RestorationAmirhosein Ghasemabadi, Muhammad Kamran Janjua, Mohammad Salameh, Di NiuNeurIPS 2024 · 被引用 28 次
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang 等AAAI 2023 · 被引用 26 次
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra 等NeurIPS 2021 · 被引用 382 次
- Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang 等ICCV 2019 · 被引用 309 次
- Omniscient Video Super-ResolutionPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang 等ICCV 2021 · 被引用 89 次
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang 等NeurIPS 2021 · 被引用 31 次
相关 Paper
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu 等CVPR 2022 · 被引用 128 次
- Video Super-Resolution With Temporal Group AttentionTakashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan 等CVPR 2020
- Trajectory-aware Shifted State Space Models for Online Video Super-ResolutionQiang Zhu, Xiandong Meng, Yuxuan Jiang, Fan Zhang 等ICLR 2026 · 被引用 3 次
- LDIP: Long Distance Information Propagation for Video Super-ResolutionMichael Bernasconi, Abdelaziz Djelouah, Yang Zhang, Markus Gross 等ICCV 2025 · 被引用 1 次
