Lune

NeurIPS2025顶会

ZeroPatcher: Training-free Sampler for Video Inpainting and Editing

Shaoshu Yang, Yingya Zhang, Ran He

2025年份
2被引次数
1顶会引用

摘要

Video inpainting and editing have long been challenging tasks in the video generation community, requiring extensive computational resources and large datasets to train models with satisfactory performance. Recent breakthroughs in large-scale video foundation models have greatly enhanced text-to-video generation capabilities. This naturally leads to the idea of leveraging the prior knowledge from these powerful generators to facilitate video inpainting and editing. In this work, we investigate the feasibility of employing pre-trained text-to-video foundation models for high-quality video inpainting and editing without additional training. Specifically, we introduce a model-agnostic denoising sampler that optimizes the trajectory by maximizing the log-likelihood expectation conditioned on the known video segments. To enable efficient dynamic object removal and replacement, we propose a latent mask fuser that performs accurate video masking directly in latent space, eliminating the need for explicit VAE decoding and encoding. We implement our approach in widely-used foundation generators such as CogVideoX and HunyuanVideo, demonstrating the model-agnostic nature of our sampler. Comprehensive quantitative and qualitative evaluations confirm that our method achieves outstanding video inpainting and editing performance in a plug-and-play fashion.

Given an input video and a dynamic mask, video inpainting or editing tasks require models to render the masked regions according to user specifications. As a long-standing challenge in video research, 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

this problem has been approached through various paradigms, including transformers (44; 17) and diffusion models (42; 15). Recent advances in foundational video generators (33; 38; 28), particularly through large diffusion transformers, have significantly improved video generation capabilities. This progress naturally suggests leveraging these powerful foundation models to advance video inpainting and editing. However, effectively utilizing their conditional generation abilities for these tasks would typically demand substantial computational resources for training, given their massive scale. Furthermore, as foundation models continue to evolve, traditional approaches relying on extensive fine-tuning will face increasing challenges in adapting to new video generators. An alternative solution is to employ these video generators as data priors, enabling task resolution in a training-free manner.

Recent research has extensively explored methods for enabling conditional generation in image diffusion models. For inpainting tasks where significant portions of the image are unavailable, current approaches typically employ back-projection operations (32) to incorporate information from known regions. However, this process does not directly estimate the conditional denoising distribution and may fail when the generation trajectory substantially deviates from the known data (32).

In this work, we introduce ZeroPatcher, a novel training-free approach that unlocks video inpainting capabilities in text-to-video foundation models. Our method is theoretically model-agnostic and compatible with diffusion-based video generators. To approximate the conditional denoising distribution, we propose conditional denoising expectation maximization (CD-EM), which formulates an EM problem over the denoising trajectory conditioned on known data. The framework consists of an expectation step implemented through Monte-Carlo sampling and a maximization step solved via fixed-point iteration. We further establish the uniqueness of the fixed point in the maximization step through theoretical analysis. While existing methods inject known information by pixel replacement at each denoising step, this approach is incompatible with prevalent latent diffusion models that cannot perform precise masking in latent space. To overcome this limitation, we present a mask fuser -a lightweight convolutional architecture that enables accurate dynamic masking in latent space.

We perform extensive evaluations on the DAVIS (22) and YouTube-VOS (35) datasets to assess both the inpainting and editing capabilities of our method. By leveraging the generative power of modern video foundation models, our approach achieves performance competitive with trained inpainting models. Furthermore, we demonstrate ZeroPatcher's video editing potential, showing superior ability to modify object shapes while maintaining surrounding content compared to existing methods. Our key contributions are:

• We introduce ZeroPatcher, a training-free framework that adapts video diffusion models for video inpainting and editing. The proposed latent mask fuser enables dynamic video masking directly in latent space.

• We present conditional denoising expectation maximization (CD-EM) to optimize sampling trajectories using known video regions as guidance, accompanied by theoretic

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖