Video Frame Interpolation Transformer
Zhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen, Ming-Hsuan Yang
摘要
Existing methods for video interpolation heavily rely on deep convolution neural networks, and thus suffer from their intrinsic limitations, such as content-agnostic kernel weights and restricted receptive field. To address these issues, we propose a Transformer-based video interpolation framework that allows content-aware aggregation weights and considers long-range dependencies with the self-attention operations. To avoid the high computational cost of global self-attention, we introduce the concept of local attention into video interpolation and extend it to the spatial-temporal domain. Furthermore, we propose a space-time separation strategy to save memory usage, which also improves performance. In addition, we develop a multi-scale frame synthesis scheme to fully realize the potential of Transformers. Extensive experiments demonstrate the proposed model performs favorably against the stateof-the-art methods both quantitatively and qualitatively on a variety of benchmark datasets. The code and models are released at https://github.com/zhshi0816/ Video-Frame-Interpolation-Transformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- AccFlow: Backward Accumulation for Long-Range Optical FlowGuangyang Wu, Xiaohong Liu, Kunming Luo, Xi Liu 等ICCV 2023 · 被引用 33 次
- Light-VQA: A Multi-Dimensional Quality Assessment Model for Low-Light Video EnhancementYunlong Dong, Xiaohong Liu, Yixuan Gao, Xunchu Zhou 等ACM MM 2023 · 被引用 23 次
- FastLLVE: Real-Time Low-Light Video Enhancement with Intensity-Aware Look-Up TableWenhao Li, Guangyang Wu, Wenyi Wang, Peiran Ren 等ACM MM 2023 · 被引用 21 次
- Sparse Global Matching for Video Frame Interpolation with Large MotionChunxu Liu, Guozhen Zhang, Rui Zhao, Limin WangCVPR 2024 · 被引用 17 次
- Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter CorrectionDelin Qu, Yizhen Lao, Zhigang Wang, Dong Wang 等ICCV 2023 · 被引用 11 次
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
相关 Paper
- Video Frame Interpolation with Flow TransformerPan Gao, Haoyue Tian, Jie QinACM MM 2023 · 被引用 4 次
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu 等CVPR 2022 · 被引用 128 次
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 等ICCV 2021 · 被引用 38 次
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai 等ICCV 2021 · 被引用 224 次
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 被引用 113 次
