Towards An End-to-End Framework for Flow-Guided Video Inpainting
Zhen Li, Chengze Lu, Jianhua Qin, Chun-Le Guo, Ming-Ming Cheng
Abstract
Optical flow, which captures motion information across frames, is exploited in recent video inpainting methods through propagating pixels along its trajectories. However, the hand-crafted flow-based processes in these methods are applied separately to form the whole inpainting pipeline. Thus, these methods are less efficient and rely heavily on the intermediate results from earlier stages. In this paper, we propose an End-to-End framework for Flow-Guided Video Inpainting (E <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> FGVI) through elaborately designed three trainable modules, namely, flow completion, feature propagation, and content hallucination modules. The three modules correspond with the three stages of previous flow-based methods but can be Jointly optimized, leading to a more efficient and effective inpainting process. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods both qualitatively and quantitatively and shows promising efficiency. The code is available at https://github.com/MCG-NKU/E2FGVI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7802b1a7-0a0d-49a9-8ad4-51cf63d92da2Cited by top-tier papers55
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- GRACE: Loss-Resilient Real-Time Video through Neural CodecsYihua Cheng, Ziyi Zhang, Hanchen Li, Anton Arapin et al.NSDI 2024 · 53 citations
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu et al.ACM MM 2023 · 42 citations
- MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D EditingChenjie Cao, Chaohui Yu, Fan Wang, Xiangyang Xue et al.NeurIPS 2024 · 36 citations
Builds on31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Video Diffusion Models Are Strong Video InpainterMinhyeok Lee, Suhwan Cho, Chajin Shin, Jungho Lee et al.AAAI 2025 · 26 citations
- Progressive Temporal Feature Alignment Network for Video InpaintingXueyan Zou, Linjie Yang, Ding Liu, Yong Jae LeeCVPR 2021
- Semi-Supervised Video Inpainting with Cycle Consistency ConstraintsZhiliang Wu, Hanyu Xuan, Changchang Sun, Weili Guan et al.CVPR 2023
- An Internal Learning Approach to Video InpaintingHaotian Zhang, Long Mai, Hailin Jin, Zhaowen Wang et al.ICCV 2019 · 77 citations
- BVINet: Unlocking Blind Video Inpainting With Zero AnnotationsZhiliang Wu, Kerui Chen, Kun Li, Hehe Fan et al.ICCV 2025 · 30 citations
