Diffusion Self-Guidance for Controllable Image Generation
Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, Aleksander Holynski
Abstract
Large-scale generative models are capable of producing high-quality images from detailed text descriptions. However, many aspects of an image are difficult or impossible to convey through text. We introduce self-guidance, a method that provides greater control over generated images by guiding the internal representations of diffusion models. We demonstrate that properties such as the shape, location, and appearance of objects can be extracted from these representations and used to steer the sampling process. Self-guidance operates similarly to standard classifier guidance, but uses signals present in the pretrained model itself, requiring no additional models or training. We show how a simple set of properties can be composed to perform challenging image manipulations, such as modifying the position or size of specific objects, merging the appearance of objects in one image with the layout of another, composing objects from multiple images into one, and more. We also show that self-guidance can be used for editing real images. See our project page for results and an interactive demo:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cf43cde-a8dd-4d57-a7f0-4b126a982ee3Cited by top-tier papers167
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan et al.ICLR 2024 · 223 citations
- DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image EditingYujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan et al.CVPR 2024 · 117 citations
- LLM-grounded Video Diffusion ModelsLong Lian, Baifeng Shi, Adam Yala, Trevor Darrell et al.ICLR 2024 · 87 citations
- OBJECT 3DIT: Language-guided 3D-aware Image EditingOscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna et al.NeurIPS 2023 · 79 citations
- Cross-Image Attention for Zero-Shot Appearance TransferYuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor et al.SIGGRAPH 2024 · 72 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion EstimatorsDaniel Geng, Andrew OwensICLR 2024 · 46 citations
- Self-Guided Diffusion ModelsVincent Tao Hu, David W. Zhang, Yuki M. Asano, Gertjan J. Burghouts et al.CVPR 2023
- Paint by Example: Exemplar-based Image Editing with Diffusion ModelsBinxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang et al.CVPR 2023
- PAIR Diffusion: A Comprehensive Multimodal Object-Level Image EditorVidit Goel, Elia Peruzzo, Yifan Jiang, Dejia Xu et al.CVPR 2024
- Continuous Control of Editing Models via Adaptive-Origin GuidanceAlon Wolf, Chen Katzir, Kfir Aberman, Or PatashnikSIGGRAPH 2026
