Retrieval-Augmented Diffusion Models
Andreas Blattmann, Robin Rombach, Kaan Oktay, Jonas Müller, Björn Ommer
Abstract
Novel architectures have recently improved generative image synthesis leading to excellent visual quality in various tasks. Much of this success is due to the scalability of these architectures and hence caused by a dramatic increase in model complexity and in the computational resources invested in training these models. Our work questions the underlying paradigm of compressing large training data into ever growing parametric representations. We rather present an orthogonal, semiparametric approach. We complement comparably small diffusion or autoregressive models with a separate image database and a retrieval strategy. During training we retrieve a set of nearest neighbors from this external database for each training instance and condition the generative model on these informative samples. While the retrieval approach is providing the (local) content, the model is focusing on learning the composition of scenes based on this content. As demonstrated by our experiments, simply swapping the database for one with different contents transfers a trained model post-hoc to a novel domain. The evaluation shows competitive performance on tasks which the generative model has not been trained on, such as class-conditional synthesis, zero-shot stylization or text-to-image synthesis without requiring paired text-image data. With negligible memory and computational overhead for the external database and retrieval we can significantly reduce the parameter count of the generative model and still outperform the state-of-the-art. * The first two authors contributed equally to this work. 36th Conference on Neural Information Processing Systems (NeurIPS 2022). 'A purple salamander in the grass.' 'A zebra-skinned panda.' 'A teddy bear riding a motorcycle.' 'Image of a monkey with the fur of a leopard.' 'A stag.' 'A basket full of fruits.' 'A woman playing piano.' 'A table set.'
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0623a696-3116-42b8-bbc3-1a63d1b0e30eCited by top-tier papers83
- Fast and Scalable Analytical DiffusionXinyi Shang, Peng Sun, Jingyu Lin, Zhiqiang ShenICML 2026 · 1,092 citations
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai et al.ICCV 2023 · 301 citations
- StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion ModelsZhizhong Wang, Lei Zhao, Wei XingICCV 2023 · 219 citations
- HyperDiffusion: Generating Implicit Neural Fields with Weight-Space DiffusionZiya Erkoç, Fangchang Ma, Qi Shan, Matthias Nießner et al.ICCV 2023 · 174 citations
- Retrieval-Augmented Multimodal Language ModelingMichihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James et al.ICML 2023 · 153 citations
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- kNN-Diffusion: Image Generation via Large-Scale RetrievalShelly Sheynin, Oron Ashual, Adam Polyak, Uriel Singer et al.ICLR 2023 · 46 citations
- Composer: Creative and Controllable Image Synthesis with Composable ConditionsLianghua Huang, Di Chen, Yu Liu, Yujun Shen et al.ICML 2023 · 371 citations
- Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion ModelsPablo Pernias, Dominic Rampas, Mats Leon Richter, Christopher Pal et al.ICLR 2024 · 60 citations
- Cross-Image Attention for Zero-Shot Appearance TransferYuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor et al.SIGGRAPH 2024 · 72 citations
- TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal ModelsLeigang Qu, Haochuan Li, Tan Wang, Wenjie Wang et al.ICLR 2025
