Re-Imagen: Retrieval-Augmented Text-to-Image Generator
Wenhu Chen, Hexiang Hu, Chitwan Saharia, William W. Cohen
摘要
Research on text-to-image generation has witnessed significant progress in generating diverse and photo-realistic images, driven by diffusion and auto-regressive models trained on large-scale image-text data. Though state-of-the-art models can generate high-quality images of common entities, they often have difficulty generating images of uncommon entities, such as Chortai (dog)' or Picarones (food)'. To tackle this issue, we present the Retrieval-Augmented Text-to-Image Generator (Re-Imagen), a generative model that uses retrieved information to produce high-fidelity and faithful images, even for rare or unseen entities. Given a text prompt, Re-Imagen accesses an external multi-modal knowledge base to retrieve relevant (image, text) pairs and uses them as references to generate the image. With this retrieval step, Re-Imagen is augmented with the knowledge of high-level semantics and low-level visual details of the mentioned entities, and thus improves its accuracy in generating the entities' visual appearances. We train Re-Imagen on a constructed dataset containing (image, text, retrieval) triples to teach the model to ground on both text prompt and retrieval. Furthermore, we develop a new sampling strategy to interleave the classifier-free guidance for text and retrieval conditions to balance the text and retrieval alignment. Re-Imagen achieves significant gain on FID score over COCO and WikiImage. To further evaluate the capabilities of the model, we introduce EntityDrawBench, a new benchmark that evaluates image generation for diverse entities, from frequent to rare, across multiple object categories including dogs, foods, landmarks, birds, and characters. Human evaluation on EntityDrawBench shows that Re-Imagen can significantly improve the fidelity of generated images, especially on less frequent entities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper94
- BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and EditingDongxu Li, Junnan Li, Steven C. H. HoiNeurIPS 2023 · 被引用 587 次
- DIRE for Diffusion-Generated Image DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang 等ICCV 2023 · 被引用 479 次
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 被引用 403 次
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai 等ICCV 2023 · 被引用 301 次
- Subject-driven Text-to-Image Generation via Apprenticeship LearningWenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz 等NeurIPS 2023 · 被引用 265 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal ModelsLeigang Qu, Haochuan Li, Tan Wang, Wenjie Wang 等ICLR 2025
- PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance PredictionEduard Gabriel Poesina, Adriana Valentina Costache, Adrian-Gabriel Chifu, Josiane Mothe 等CVPR 2025
- Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM GuidanceDongmin Park, Sebin Kim, Taehong Moon, Minkyu Kim 等ICLR 2025
- Retrieval-Augmented Multimodal Language ModelingMichihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James 等ICML 2023 · 被引用 153 次
