ZeroPatcher: Training-free Sampler for Video Inpainting and Editing
Shaoshu Yang, Yingya Zhang, Ran He
摘要
Video inpainting and editing have long been challenging tasks in the video generation community, requiring extensive computational resources and large datasets to train models with satisfactory performance. Recent breakthroughs in large-scale video foundation models have greatly enhanced text-to-video generation capabilities. This naturally leads to the idea of leveraging the prior knowledge from these powerful generators to facilitate video inpainting and editing. In this work, we investigate the feasibility of employing pre-trained text-to-video foundation models for high-quality video inpainting and editing without additional training. Specifically, we introduce a model-agnostic denoising sampler that optimizes the trajectory by maximizing the log-likelihood expectation conditioned on the known video segments. To enable efficient dynamic object removal and replacement, we propose a latent mask fuser that performs accurate video masking directly in latent space, eliminating the need for explicit VAE decoding and encoding. We implement our approach in widely-used foundation generators such as CogVideoX and HunyuanVideo, demonstrating the model-agnostic nature of our sampler. Comprehensive quantitative and qualitative evaluations confirm that our method achieves outstanding video inpainting and editing performance in a plug-and-play fashion.
Given an input video and a dynamic mask, video inpainting or editing tasks require models to render the masked regions according to user specifications. As a long-standing challenge in video research, 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
this problem has been approached through various paradigms, including transformers (44; 17) and diffusion models (42; 15). Recent advances in foundational video generators (33; 38; 28), particularly through large diffusion transformers, have significantly improved video generation capabilities. This progress naturally suggests leveraging these powerful foundation models to advance video inpainting and editing. However, effectively utilizing their conditional generation abilities for these tasks would typically demand substantial computational resources for training, given their massive scale. Furthermore, as foundation models continue to evolve, traditional approaches relying on extensive fine-tuning will face increasing challenges in adapting to new video generators. An alternative solution is to employ these video generators as data priors, enabling task resolution in a training-free manner.
Recent research has extensively explored methods for enabling conditional generation in image diffusion models. For inpainting tasks where significant portions of the image are unavailable, current approaches typically employ back-projection operations (32) to incorporate information from known regions. However, this process does not directly estimate the conditional denoising distribution and may fail when the generation trajectory substantially deviates from the known data (32).
In this work, we introduce ZeroPatcher, a novel training-free approach that unlocks video inpainting capabilities in text-to-video foundation models. Our method is theoretically model-agnostic and compatible with diffusion-based video generators. To approximate the conditional denoising distribution, we propose conditional denoising expectation maximization (CD-EM), which formulates an EM problem over the denoising trajectory conditioned on known data. The framework consists of an expectation step implemented through Monte-Carlo sampling and a maximization step solved via fixed-point iteration. We further establish the uniqueness of the fixed point in the maximization step through theoretical analysis. While existing methods inject known information by pixel replacement at each denoising step, this approach is incompatible with prevalent latent diffusion models that cannot perform precise masking in latent space. To overcome this limitation, we present a mask fuser -a lightweight convolutional architecture that enables accurate dynamic masking in latent space.
We perform extensive evaluations on the DAVIS (22) and YouTube-VOS (35) datasets to assess both the inpainting and editing capabilities of our method. By leveraging the generative power of modern video foundation models, our approach achieves performance competitive with trained inpainting models. Furthermore, we demonstrate ZeroPatcher's video editing potential, showing superior ability to modify object shapes while maintaining surrounding content compared to existing methods. Our key contributions are:
• We introduce ZeroPatcher, a training-free framework that adapts video diffusion models for video inpainting and editing. The proposed latent mask fuser enables dynamic video masking directly in latent space.
• We present conditional denoising expectation maximization (CD-EM) to optimize sampling trajectories using known video regions as guidance, accompanied by theoretic
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 被引用 370 次
- FateZero: Fusing Attentions for Zero-shot Text-based Video EditingChenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei 等ICCV 2023 · 被引用 510 次
- Fuse Your Latents: Video Editing with Multi-source Latent Diffusion ModelsTianyi Lu, Xing Zhang, Jiaxi Gu, Renjing Pei 等ACM MM 2024 · 被引用 2 次
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi 等KDD 2026
- MaskINT: Video Editing via Interpolative Non-autoregressive Masked TransformersHaoyu Ma, Shahin Mahdizadehaghdam, Bichen Wu, Zhipeng Fan 等CVPR 2024 · 被引用 3 次
