Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis
Nithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang, Toshiaki Koike-Akino, Vishal M. Patel, Tim K. Marks
摘要
Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained on the generative task, to guide the generative process (e.g., using language). However, such guidance is typically useful only towards synthesizing high-level semantics rather than editing fine-grained details as in image-to-image translation tasks. To this end, and capitalizing on the powerful fine-grained generative control offered by the recent diffusion-based generative * Work done during internship at MERL. models, we introduce Steered Diffusion, a generalized framework for photorealistic zero-shot conditional image generation using a diffusion model trained for unconditional generation. The key idea is to steer the image generation of the diffusion model at inference time via designing a loss using a pre-trained inverse model that characterizes the conditional task. This loss modulates the sampling trajectory of the diffusion process. Our framework allows for easy incorporation of multiple conditions during inference. We present experiments using steered diffusion on several tasks including inpainting, colorization, text-guided semantic editing, and image super-resolution. Our results demonstrate clear qualitative and quantitative improvements over state-of-the-art diffusion-based plug-and-play models while adding negligible additional computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- TFG: Unified Training-Free Guidance for Diffusion ModelsHaotian Ye, Haowei Lin, Jiaqi Han, Minkai Xu 等NeurIPS 2024 · 被引用 118 次
- TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsHaomiao Ni, Bernhard Egger, Suhas Lohit, Anoop Cherian 等CVPR 2024 · 被引用 7 次
- ZePo: Zero-Shot Portrait Stylization with Faster SamplingJin Liu, Huaibo Huang, Jie Cao, Ran HeACM MM 2024 · 被引用 6 次
- Zero-Shot Conditioning of Score-Based Diffusion Models by Neuro-Symbolic ConstraintsDavide Scassola, Sebastiano Saccani, Ginevra Carbone, Luca BortolussiAAAI 2025 · 被引用 2 次
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image DetectorsSepehr Dehdashtian, Mashrur Mahmud Morshed, Jacob H. Seidman, Gaurav Bharaj 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
- Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style TransferSerin Yang, Hyunmin Hwang, Jong Chul YeICCV 2023 · 被引用 94 次
- Training-Free Reward-Guided Image Editing via Trajectory Optimal ControlJinho Chang, Jaemin Kim, Jong Chul YeICLR 2026 · 被引用 2 次
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
- Loss-Guided Diffusion Models for Plug-and-Play Controllable GenerationJiaming Song, Qinsheng Zhang, Hongxu Yin, Morteza Mardani 等ICML 2023 · 被引用 222 次
