Boundary Guided Learning-Free Semantic Control with Diffusion Models
Ye Zhu, Yu Wu, Zhiwei Deng, Olga Russakovsky, Yan Yan
摘要
Applying pre-trained generative denoising diffusion models (DDMs) for downstream tasks such as image semantic editing usually requires either fine-tuning DDMs or learning auxiliary editing networks in the existing literature. In this work, we present our BoundaryDiffusion method for efficient, effective and lightweight semantic control with frozen pre-trained DDMs, without learning any extra networks. As one of the first learning-free diffusion editing works, we start by seeking a comprehensive understanding of the intermediate high-dimensional latent spaces by theoretically and empirically analyzing their probabilistic and geometric behaviors in the Markov chain. We then propose to further explore the critical step for editing in the denoising trajectory that characterizes the convergence of a pre-trained DDM and introduce an automatic search method. Last but not least, in contrast to the conventional understanding that DDMs have relatively poor semantic behaviors, we prove that the critical latent space we found already exhibits semantic subspace boundaries at the generic level in unconditional DDMs, which allows us to do controllable manipulation by guiding the denoising trajectory towards the targeted boundary via a single-step operation. We conduct extensive experiments on multiple DPMs architectures (DDPM, iDDPM) and datasets (CelebA, CelebA-HQ, LSUN-church, LSUN-bedroom, AFHQ-dog) with different resolutions (64, 256), achieving superior or state-of-the-art performance in various task scenarios (image semantic editing, text-based editing, unconditional semantic control) to demonstrate the effectiveness. Project page at https://l-yezhu.github.io/BoundaryDiffusion/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial OptimizationHuayang Huang, Yu Wu, Qian WangNeurIPS 2024 · 被引用 73 次
- Interpreting the Weight Space of Customized Diffusion ModelsAmil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal 等NeurIPS 2024 · 被引用 40 次
- Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image EditingSiyi Chen, Huijie Zhang, Minzhe Guo, Yifu Lu 等NeurIPS 2024 · 被引用 29 次
- Novel Object Synthesis via Adaptive Text-Image HarmonyZeren Xiong, Zedong Zhang, Zikun Chen, Shuo Chen 等NeurIPS 2024 · 被引用 15 次
- All-in-One Slider for Attribute Manipulation in Diffusion ModelsWeixin Ye, Hongguang Zhu, Wei Wang, Yahui Liu 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Unsupervised Region-Based Image Editing of Denoising Diffusion ModelsZixiang Li, Yue Song, Renshuai Tao, Xiaohong Jia 等AAAI 2025 · 被引用 1 次
- Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion ModelsRui Jiang, Xinghe Fu, Guangcong Zheng, Teng Li 等AAAI 2025 · 被引用 2 次
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion ModelsAleksandar Cvejic, Abdelrahman Eldesokey, Peter WonkaSIGGRAPH 2025 · 被引用 3 次
- PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion ModelsWonyong Seo, Jaeho Moon, Jaehyup Lee, Soo Ye Kim 等CVPR 2026 · 被引用 2 次
- Diffusion Models Already Have A Semantic Latent SpaceMingi Kwon, Jaeseok Jeong, Youngjung UhICLR 2023 · 被引用 52 次
