Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation
Guozhen Zhang, Yuhan Zhu, Haonan Wang, Youxin Chen, Gangshan Wu, Limin Wang
摘要
Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or elaborate separate modules for each type of information, which lead to representation ambiguity and low efficiency. In this paper, we propose a novel module to explicitly extract motion and appearance information via a unifying operation. Specifically, we rethink the information process in inter-frame attention and reuse its attention map for both appearance feature enhancement and motion information extraction. Furthermore, for efficient VFI, our proposed module could be seamlessly integrated into a hybrid CNN and Transformer architecture. This hybrid pipeline can alleviate the computational complexity of inter-frame attention as well as preserve detailed lowlevel structure information. Experimental results demonstrate that, for both fixed-and arbitrary-timestep interpolation, our method achieves state-of-the-art performance on various datasets. Meanwhile, our approach enjoys a lighter computation overhead over models with close performance. The source code and models are available at https://github.com/MCG-NJU/EMA-VFI .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- MGMAE: Motion Guided Masking for Video Masked AutoencodingBingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao 等ICCV 2023 · 被引用 58 次
- VFIMamba: Video Frame Interpolation with State Space ModelsGuozhen Zhang, Chunxu Liu, Yutao Cui, Xiaotong Zhao 等NeurIPS 2024 · 被引用 48 次
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- Dual DETRs for Multi-Label Temporal Action DetectionYuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu 等CVPR 2024 · 被引用 25 次
- Generalizable Implicit Motion Modeling for Video Frame InterpolationZujin Guo, Wei Li, Chen Change LoyNeurIPS 2024 · 被引用 24 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell 等NeurIPS 2021 · 被引用 974 次
相关 Paper
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu 等CVPR 2022 · 被引用 128 次
- Video Frame Interpolation with Flow TransformerPan Gao, Haoyue Tian, Jie QinACM MM 2023 · 被引用 4 次
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
- Multi-Scale Coarse-to-Fine Transformer for Frame InterpolationChen Li, Li Song, Xueyi Zou, Jiaming Guo 等ACM MM 2022 · 被引用 1 次
- Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Yijie Yu, Ling Yang, Chujun Qin 等ACM MM 2024 · 被引用 10 次
