End-to-End Diffusion Latent Optimization Improves Classifier Guidance
Bram Wallace, Akash Gokul, Stefano Ermon, Nikhil Naik
Abstract
Classifier guidance-using the gradients of an image classifier to steer the generations of a diffusion modelhas the potential to dramatically expand the creative control over image generation and editing. However, currently classifier guidance requires either training new noiseaware models to obtain accurate gradients or using a onestep denoising approximation of the final generation, which leads to misaligned gradients and sub-optimal control.We highlight this approximation's shortcomings and propose a novel guidance method: Direct Optimization of Diffusion Latents (DOODL), which enables plug-and-play guidance by optimizing diffusion latents w.r.t. the gradients of a pre-trained classifier on the true generated pixels, using an invertible diffusion process to achieve memory-efficient backpropagation. Showcasing the potential of more precise guidance, DOODL outperforms one-step classifier guidance on computational and human evaluation metrics across different forms of guidance: using CLIP guidance to improve generations of complex prompts from DrawBench, using fine-grained visual classifiers to expand the vocabulary of Stable Diffusion, enabling image-conditioned generation with a CLIP visual encoder, and improving image aesthetics using an aesthetic scoring network. Code at https://github.com/salesforce/DOODL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers59
- SVDiff: Compact Parameter Space for Diffusion Fine-TuningLigong Han, Yinxiao Li, Han Zhang, Peyman Milanfar et al.ICCV 2023 · 384 citations
- Directly Fine-Tuning Diffusion Models on Differentiable RewardsKevin Clark, Paul Vicol, Kevin Swersky, David J. FleetICLR 2024 · 377 citations
- Manifold Preserving Guided DiffusionYutong He, Naoki Murata, Chieh-Hsin Lai, Yuhta Takida et al.ICLR 2024 · 148 citations
- ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise OptimizationLuca Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy et al.NeurIPS 2024 · 131 citations
- Finetuning Text-to-Image Diffusion Models for FairnessXudong Shen, Chao Du, Tianyu Pang, Min Lin et al.ICLR 2024 · 97 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Dynamic Classifier-Free Diffusion Guidance via Online FeedbackPinelopi Papalampidi, Olivia Wiles, Ira Ktena, Aleksandar Shtedritski et al.ICLR 2026 · 12 citations
- Elucidating the design space of classifier-guided diffusion generationJiajun Ma, Tianyang Hu, Wenjia Wang, Jiacheng SunICLR 2024 · 24 citations
- SpaText: Spatio-Textual Representation for Controllable Image GenerationOmri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta et al.CVPR 2023
- Guiding a Diffusion Model by Swapping Its TokensWeijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran et al.CVPR 2026 · 2 citations
