Unsupervised Region-Based Image Editing of Denoising Diffusion Models
Zixiang Li, Yue Song, Renshuai Tao, Xiaohong Jia, Yao Zhao, Wei Wang
Abstract
Although diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper, we propose a method to identify semantic attributes in the latent space of pre-trained diffusion models without any further training. By projecting the Jacobian of the targeted semantic region into a low-dimensional subspace which is orthogonal to the non-masked regions, our approach facilitates precise semantic discovery and control over local masked areas, eliminating the need for annotations. We conducted extensive experiments across multiple datasets and various architectures of diffusion models, achieving state-of-the-art performance. In particular, for some specific face attributes, the performance of our proposed method even surpasses that of supervised approaches, demonstrating its superior ability in editing local image properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5351e9b4-9fce-4a21-9dc0-8954c2e01f87Cited by top-tier papers2
- DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image EditingZixiang Li, Haoyu Wang, Wei Wang, Chuangchuang Tan et al.NeurIPS 2025 · 5 citations
- RAIN: Redundancy-Aware Latent Injection for Quality-Preserving Image WatermarkingYehan Sun, Rongrong Ni, Chuangchuang Tan, Huan Liu et al.AAAI 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion ModelsYusuf Dalva, Pinar YanardagCVPR 2024
- Boundary Guided Learning-Free Semantic Control with Diffusion ModelsYe Zhu, Yu Wu, Zhiwei Deng, Olga Russakovsky et al.NeurIPS 2023 · 40 citations
- DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image EditingHaozhe Jia, Yan Li, Hengfei Cui, Di Xu et al.ACM MM 2024 · 2 citations
- Pixel-Perfect Puppetry: Precision-Guided Enhancement for Face Image and Video EditingYan Li, Zhenyi Wang, Guanghao Li, Wei Xue et al.ICLR 2026
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion ModelsAleksandar Cvejic, Abdelrahman Eldesokey, Peter WonkaSIGGRAPH 2025 · 3 citations
