Learning Trajectory-Aware Transformer for Video Super-Resolution
Chengxu Liu, Huan Yang, Jianlong Fu, Xueming Qian
Abstract
Video super-resolution (VSR) aims to restore a sequence of high-resolution (HR) frames from their low-resolution (LR) counterparts. Although some progress has been made, there are grand challenges to effectively utilize temporal dependency in entire video sequences. Existing approaches usually align and aggregate video frames from limited adjacent frames (e.g., 5 or 7 frames), which prevents these approaches from satisfactory results. In this paper, we take one step further to enable effective spatio-temporal learning in videos. We propose a novel Trajectory-aware Transformer for Video Super-Resolution (TTVSR). In particular, we formulate video frames into several pre-aligned trajectories which consist of continuous visual tokens. For a query token, self-attention is only learned on relevant visual tokens along spatio-temporal trajectories. Compared with vanilla vision Transformers, such a design significantly reduces the computational cost and enables Transformers to model long-range features. We further propose a cross-scale feature tokenization module to over-come scale-changing problems that often occur in long-range videos. Experimental results demonstrate the superiority of the proposed TTVSR over state-of-the-art models, by extensive quantitative and qualitative evaluations in four widely-used video super-resolution benchmarks. Both code and pre-trained models can be downloaded at https://github.com/researchmm/TTVSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae9d1afd-a21f-41c4-bf53-192ff088d5d1Cited by top-tier papers32
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan et al.NeurIPS 2022 · 318 citations
- Zero-Reference Low-Light Enhancement via Physical Quadruple PriorsWenjing Wang, Huan Yang, Jianlong Fu, Jiaying LiuCVPR 2024 · 90 citations
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li et al.CVPR 2026 · 42 citations
- Learning Truncated Causal History Model for Video RestorationAmirhosein Ghasemabadi, Muhammad Kamran Janjua, Mohammad Salameh, Di NiuNeurIPS 2024 · 28 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra et al.NeurIPS 2021 · 382 citations
- Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang et al.ICCV 2019 · 309 citations
- Omniscient Video Super-ResolutionPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang et al.ICCV 2021 · 89 citations
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang et al.NeurIPS 2021 · 31 citations
Related papers
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu et al.CVPR 2022 · 128 citations
- Video Super-Resolution With Temporal Group AttentionTakashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan et al.CVPR 2020
- Trajectory-aware Shifted State Space Models for Online Video Super-ResolutionQiang Zhu, Xiandong Meng, Yuxuan Jiang, Fan Zhang et al.ICLR 2026 · 3 citations
- LDIP: Long Distance Information Propagation for Video Super-ResolutionMichael Bernasconi, Abdelaziz Djelouah, Yang Zhang, Markus Gross et al.ICCV 2025 · 1 citation
