Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis
Nithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang, Toshiaki Koike-Akino, Vishal M. Patel, Tim K. Marks
Abstract
Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained on the generative task, to guide the generative process (e.g., using language). However, such guidance is typically useful only towards synthesizing high-level semantics rather than editing fine-grained details as in image-to-image translation tasks. To this end, and capitalizing on the powerful fine-grained generative control offered by the recent diffusion-based generative * Work done during internship at MERL. models, we introduce Steered Diffusion, a generalized framework for photorealistic zero-shot conditional image generation using a diffusion model trained for unconditional generation. The key idea is to steer the image generation of the diffusion model at inference time via designing a loss using a pre-trained inverse model that characterizes the conditional task. This loss modulates the sampling trajectory of the diffusion process. Our framework allows for easy incorporation of multiple conditions during inference. We present experiments using steered diffusion on several tasks including inpainting, colorization, text-guided semantic editing, and image super-resolution. Our results demonstrate clear qualitative and quantitative improvements over state-of-the-art diffusion-based plug-and-play models while adding negligible additional computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cb12bbb-3b09-4365-b727-79b0a0f7395eCited by top-tier papers10
- TFG: Unified Training-Free Guidance for Diffusion ModelsHaotian Ye, Haowei Lin, Jiaqi Han, Minkai Xu et al.NeurIPS 2024 · 118 citations
- TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsHaomiao Ni, Bernhard Egger, Suhas Lohit, Anoop Cherian et al.CVPR 2024 · 7 citations
- ZePo: Zero-Shot Portrait Stylization with Faster SamplingJin Liu, Huaibo Huang, Jie Cao, Ran HeACM MM 2024 · 6 citations
- Zero-Shot Conditioning of Score-Based Diffusion Models by Neuro-Symbolic ConstraintsDavide Scassola, Sebastiano Saccani, Ginevra Carbone, Luca BortolussiAAAI 2025 · 2 citations
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image DetectorsSepehr Dehdashtian, Mashrur Mahmud Morshed, Jacob H. Seidman, Gaurav Bharaj et al.NeurIPS 2025 · 1 citation
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 104 citations
- Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style TransferSerin Yang, Hyunmin Hwang, Jong Chul YeICCV 2023 · 94 citations
- Training-Free Reward-Guided Image Editing via Trajectory Optimal ControlJinho Chang, Jaemin Kim, Jong Chul YeICLR 2026 · 2 citations
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
- Loss-Guided Diffusion Models for Plug-and-Play Controllable GenerationJiaming Song, Qinsheng Zhang, Hongxu Yin, Morteza Mardani et al.ICML 2023 · 222 citations
