ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models
Peijie Qiu, Hariharan Ramshankar, Arnau Ramisa, Amit C C, Rene Vidal, Vamsi Salaka, Rahul Bhagat
摘要
Diffusion models have emerged as the leading approach for text-to-image generation. However, their iterative sampling process, which gradually morphs random noise into coherent images, introduces significant latency that limits their applicability. While recent few-step diffusion models reduce the number of sampling steps to as few as one to four steps, they often compromise image quality and prompt alignment, especially in one-step generation. Additionally, these models require computationally expensive training procedures. To address these limitations, we propose ImageRAGTurbo, a novel approach to efficiently finetune few-step diffusion models via retrieval augmentation. Given a text prompt, we retrieve relevant text-image pairs from a database and use them to condition the generation process. We argue that such retrieved examples provide rich contextual information to the UNet denoiser that helps reduce the number of denoising steps without compromising image quality. Indeed, our initial investigations show that using the retrieved content to edit the denoiser's latent space (-space) without additional finetuning already improves prompt fidelity. To further improve the quality of the generated images, we augment the UNet denoiser with a trainable adapter in the -space, which efficiently blends the retrieved content with the target prompt using a cross-attention mechanism. Experimental results on fast text-to-image generation demonstrate that our approach produces high-fidelity images without compromising latency compared to existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANsYanwu Xu, Yang Zhao, Zhisheng Xiao, Tingbo HouCVPR 2024
- DiffuseHigh: Training-Free Progressive High-Resolution Image Synthesis Through Structure GuidanceYounghyun Kim, Geunmin Hwang, Junyu Zhang, Eunbyung ParkAAAI 2025 · 被引用 30 次
- Multistep Distillation of Diffusion Models via Moment MatchingTim Salimans, Thomas Mensink, Jonathan Heek, Emiel HoogeboomNeurIPS 2024 · 被引用 93 次
- SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language ModelsShanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin 等ACM MM 2023 · 被引用 45 次
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
