Lune

ICLR2024Top-tier venue

PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, Qiang Xu

2024Year
166Citations
79Top-tier citations

Abstract

ABSTRACT Text-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusionbased editing, where a source image is edited according to a target prompt, the process commences by acquiring a noisy latent vector corresponding to the source image via the diffusion model. This vector is subsequently fed into separate source and target diffusion branches for editing. The accuracy of this inversion process significantly impacts the final editing outcome, influencing both essential content preservation of the source image and edit fidelity according to the target prompt. Prior inversion techniques aimed at finding a unified solution in both the source and target diffusion branches. However, our theoretical and empirical analyses reveal that disentangling these branches leads to a distinct separation of responsibilities for preserving essential content and ensuring edit fidelity. Building on this insight, we introduce "PnP Inversion," a novel technique achieving optimal performance of both branches with just three lines of code. To assess image editing performance, we present PIE-Bench, an editing benchmark with 700 images showcasing diverse scenes and editing types, accompanied by versatile annotations and comprehensive evaluation metrics. Compared to state-of-the-art optimization-based inversion techniques, our solution not only yields superior performance across 8 editing methods but also achieves nearly an order of speed-up. We assume a 2-step diffusion process for illustration. Due to nonexistent of ideal z I 2 , common practice uses DDIM Inversion (Song et al., 2020) to approximate z I t , resulting in z Ip t with perturbation. Diffusionbased editing methods start from the perturbed noisy latent z Ip 2 and perform DDIM sampling in a source and a target diffusion branch, further resulting in the distance shown on the figure. Null-Text Inversion and StyleDiffusion optimize a specific latent used in both source and target branches to reduce this distance. Negative-Prompt Inversion assigns the guidance scale to 1 to decrease the distance. In contrast, PnP Inversion disentangles source and target branches in editing. By leaving the target diffusion branch untouched, PnP Inversion retains the edit fidelity. By directly returning the source branch to z src 0 , PnP Inversion achieves the best possible essential content preservation. We use numbers to mark operation step order, where solid circles are steps added by PnP Inversion. Text-guided diffusion models (Rombach et al., 2022; Ramesh et al., 2022) have become the mainstream image generation technique, praised for their realism and diversity. As the noise latent space of diffusion models (Meng et al., 2022; Kawar et al., 2023; Hertz et al., 2023; Balaji et al., 2022; Liu et al., 2023a) possesses the capacity to retain and modify images, we can perform prompt-based editing with diffusion models, where a source image is edited according to a target prompt. The common practice is to maintain two diffusion branches: one for the source image and the other for the target image. By carefully exchanging information between the two branches, we can preserve the essential content in the source image while achieving edit fidelity according to the target prompt. However, such manipulations are only straightforward when the diffusion latent space (noisy latent in each diffusion step) corresponding to the source image is available. When editing images without known latent space, we have to invert the diffusion model to obtain their latent vectors first. While DDIM inversion is effective for unconditional diffusion (Song et al., 2020; Dhariwal & Nichol, 2021) , much of the research (Hertz et al., 2023; Han et al., 2023) has centered on inverting the diffusion process with conditional inputs. This is driven by the significance of conditions in applications like text-based image editing. However, introducing conditions undermines DDIM inversion quality, as evidenced in Figure 2 . With the advent of Null-Text Inversion (Mokady et al., 2023) , a prevailing consensus (Dong et al., 2023; Li et al., 2023b) has emerged: achieving superior inversion foot_0 necessitates rigorous optimization. Methods that forgo such optimization, such as Negative-Prompt Inversion (Miyake et al., 2023) , compromise editing outcomes. In this paper, we challenge this prevailing wisdom, posing two fundamental questions: What exactly are these optimization-based inversion methods truly aiming at? And, are such optimizations indispensable for diffusion-based image editing? As illustrated in Figure 2 , prior optimization-based approaches strive to minimize the distance between z src 0 and z Fsg 0 /z Ftg 0 by indirectly tweaking the generation model's input parameters. Given the magnitude of the optimization network, like UNet, and the impracticality of prolonged optimization durations, these methods often optimize the

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers79

Ask how each one uses it

Builds on41

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines