High-Fidelity Guided Image Synthesis with Latent Diffusion Models
Jaskirat Singh, Stephen Gould, Liang Zheng
摘要
Controllable image synthesis with user scribbles has gained huge public interest with the recent advent of textconditioned latent diffusion models. The user scribbles control the color composition while the text prompt provides control over the overall image semantics. However, we note that prior works in this direction suffer from an intrinsic domain shift problem wherein the generated outputs often lack details and resemble simplistic representations of the target domain. In this paper, we propose a novel guided image synthesis framework, which addresses this problem by modelling the output image as the solution of a constrained optimization problem. We show that while computing an exact solution to the optimization is infeasible, an approximation of the same can be achieved while just requiring a single pass of the reverse diffusion process. Additionally, we show that by simply defining a cross-attention based correspondence between the input text tokens and the user stroke-painting, the user is also able to control the semantics of different painted regions without requiring any conditional training or finetuning. Human user study results show that the proposed approach outperforms the previous state-of-the-art by over 85.32% on the overall user satisfaction scores. Project page for our paper is available at https://1jsingh.github.io/gradop .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan 等ICLR 2024 · 被引用 223 次
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui 等ICCV 2023 · 被引用 55 次
- Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA FeedbackJaskirat Singh, Liang ZhengNeurIPS 2023 · 被引用 48 次
- Towards Language-Driven Video Inpainting via Multimodal Large Language ModelsJianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou 等CVPR 2024 · 被引用 20 次
- Genuine Knowledge from Practice: Diffusion Test-Time Adaptation for Video Adverse Weather RemovalYijun Yang, Hongtao Wu, Angelica I. Avilés-Rivero, Yulun Zhang 等CVPR 2024 · 被引用 18 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie 等CVPR 2024 · 被引用 26 次
- Local Conditional Controlling for Text-to-Image Diffusion ModelsYibo Zhao, Liang Peng, Yang Yang, Zekai Luo 等AAAI 2025 · 被引用 2 次
- Guided Image-to-Image Translation With Bi-Directional Feature TransformationBadour Albahar, Jia-Bin HuangICCV 2019 · 被引用 102 次
- Zero-shot spatial layout conditioning for text-to-image diffusion modelsGuillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière 等ICCV 2023 · 被引用 82 次
- Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationXiang Gao, Zhengbo Xu, Junhan Zhao, Jiaying LiuAAAI 2024 · 被引用 23 次
