Video Diffusion Models Are Strong Video Inpainter
Minhyeok Lee, Suhwan Cho, Chajin Shin, Jungho Lee, Sunghun Yang, Sangyoun Lee
Abstract
Propagation-based video inpainting using optical flow at the pixel or feature level has recently garnered significant attention. However, it has limitations such as the inaccuracy of optical flow prediction and the propagation of noise over time. These issues result in non-uniform noise and time consistency problems throughout the video, which are particularly pronounced when the removed area is large and involves substantial movement. To address these issues, we propose a novel First Frame Filling Video Diffusion Inpainting model (FFF-VDI). We design FFF-VDI inspired by the capabilities of pre-trained image-to-video diffusion models that can transform the first frame image into a highly natural video. To apply this to the video inpainting task, we propagate the noise latent information of future frames to fill the masked areas of the first frame's noise latent code. Next, we fine-tune the pre-trained image-to-video diffusion model to generate the inpainted video. The proposed model addresses the limitations of existing methods that rely on optical flow quality, producing much more natural and temporally consistent videos. This proposed approach is the first to effectively integrate image-to-video diffusion models into video inpainting tasks. Through various comparative experiments, we demonstrate that the proposed model can robustly handle diverse inpainting types with high quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1effc166-60a3-4d47-b838-be407580e864Cited by top-tier papers13
- MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalBojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang et al.NeurIPS 2025 · 43 citations
- ROSE: Remove Objects with Side Effects in VideosChenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao et al.NeurIPS 2025 · 37 citations
- Unified In-Context Video EditingZixuan Ye, Xuanhua He, Quande Liu, Qiulin Wang et al.ICLR 2026 · 37 citations
- Refaçade: Editing Object with Given Reference TextureYouze Huang, Penghui Ruan, Bojia Zi, Xianbiao Qi et al.CVPR 2026 · 3 citations
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video InpaintingJiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma et al.ICCV 2025 · 2 citations
Builds on9
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingRui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi et al.ICCV 2021 · 165 citations
- Towards An End-to-End Framework for Flow-Guided Video InpaintingZhen Li, Chengze Lu, Jianhua Qin, Chun-Le Guo et al.CVPR 2022 · 136 citations
Related papers
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang et al.ACM MM 2023 · 12 citations
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
- AVID: Any-Length Video Inpainting with Diffusion ModelZhixing Zhang, Bichen Wu, Xiaoyan Wang, Yaqiao Luo et al.CVPR 2024 · 25 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li et al.ACM MM 2025 · 2 citations
