VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
Yuxuan Bian, Zhaoyang Zhang, Xuan Ju, Mingdeng Cao, Liangbin Xie, Ying Shan, Qiang Xu
摘要
Video inpainting, crucial for the media industry, aims to restore corrupted content. However, current methods relying on limited pixel propagation or single-branch image inpainting architectures face challenges with generating fully masked objects, balancing background preservation with foreground generation, and maintaining ID consistency over long video. To address these issues, we propose VideoPainter, an efficient dual-branch framework featuring a lightweight context encoder. This plug-and-play encoder processes masked videos and injects background guidance into any pre-trained video diffusion transformer, generalizing across arbitrary mask types, enhancing background integration and foreground generation, and enabling user-customized control. We further introduce a strategy to resample inpainting regions for maintaining ID consistency in any-length video inpainting. Additionally, we develop a scalable dataset pipeline using advanced vision models and construct VPData and VPBench—the largest video inpainting dataset with segmentation masks and dense caption (>390K clips) —to support large-scale training and evaluation. We also show VideoPainter’s promising potential in downstream applications such as video editing. Extensive experiments demonstrate VideoPainter’s state-of-the-art performance in any-length video inpainting and editing across 8 key metrics, including video quality, mask region preservation, and textual coherence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalBojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang 等NeurIPS 2025 · 被引用 43 次
- Omni-Effects: Unified and Spatially-Controllable Visual Effects GenerationFangyuan Mao, Aiming Hao, Jintao Chen, Dongxia Liu 等AAAI 2026 · 被引用 20 次
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward OptimizationXiaoyan Cong, Haotian Yang, Angtian Wang, Yizhi Wang 等CVPR 2026 · 被引用 16 次
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect ErasingYANG FU, Yike Zheng, Ziyun Dai, Henghui DingCVPR 2026 · 被引用 14 次
- Video-As-Prompt: Unified Semantic Control for Video GenerationYuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi 等ICLR 2026 · 被引用 13 次
它引用的顶会 Paper24
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- AVID: Any-Length Video Inpainting with Diffusion ModelZhixing Zhang, Bichen Wu, Xiaoyan Wang, Yaqiao Luo 等CVPR 2024 · 被引用 25 次
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
- Object-WIPER: Training-Free Object and Associated Effect Removal in VideosSaksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep KulkarniCVPR 2026 · 被引用 5 次
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video GenerationMingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing 等CVPR 2026 · 被引用 8 次
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang 等ACM MM 2023 · 被引用 12 次
