Video Frame Interpolation with Transformer
Liying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu, Jiaya Jia
Abstract
Video frame interpolation (VFI), which aims to synthesize intermediate frames of a video, has made remarkable progress with development of deep convolutional networks over past years. Existing methods built upon convolutional networks generally face challenges of handling large motion due to the locality of convolution operations. To overcome this limitation, we introduce a novel framework, which takes advantage of Transformer to model long-range pixel correlation among video frames. Further, our network is equipped with a novel cross-scale window-based attention mechanism, where cross-scale windows interact with each other. This design effectively enlarges the receptive field and aggregates multi-scale information. Extensive quantitative and qualitative experiments demonstrate that our method achieves new state-of-the-art results on various benchmarks. The source code is available at https://github.com/dvlab-research/VFIformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers55
- VFIMamba: Video Frame Interpolation with State Space ModelsGuozhen Zhang, Chunxu Liu, Yutao Cui, Xiaotong Zhao et al.NeurIPS 2024 · 48 citations
- Generalizable Implicit Motion Modeling for Video Frame InterpolationZujin Guo, Wei Li, Chen Change LoyNeurIPS 2024 · 24 citations
- DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic CodesHao Yan, Zhihui Ke, Xiaobo Zhou, Tie Qiu et al.CVPR 2024 · 18 citations
- PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementZhanfeng Feng, Long Peng, Xin Di, Yong Guo et al.NeurIPS 2025 · 17 citations
- Sparse Global Matching for Video Frame Interpolation with Large MotionChunxu Liu, Guozhen Zhang, Rui Zhao, Limin WangCVPR 2024 · 17 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
Related papers
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- Video Frame Interpolation with Flow TransformerPan Gao, Haoyue Tian, Jie QinACM MM 2023 · 4 citations
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 113 citations
- AMT: All-Pairs Multi-Field Transforms for Efficient Frame InterpolationZhen Li, Zuo-Liang Zhu, Linghao Han, Qibin Hou et al.CVPR 2023
- Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame InterpolationGuozhen Zhang, Yuhan Zhu, Haonan Wang, Youxin Chen et al.CVPR 2023
