RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
Yuanhuiyi Lyu, Xu Zheng, Lutao Jiang, Yibo Yan, Xin Zou, Huiyu Zhou, Linfeng Zhang, Xuming Hu
摘要
Recent text-to-image generative models, e.g., Stable Diffusion V3 and Flux, have achieved notable progress. However, these models are strongly restricted to their limited knowledge, a.k.a., their own fixed parameters, that are trained with closed datasets. This leads to significant hallucinations or distortions when facing fine-grained and unseen novel real-world objects, e.g., the appearance of the Tesla Cybertruck. To this end, we present the first real-object-based retrieval-augmented generation framework (RealRAG), which augments fine-grained and unseen novel object generation by learning and retrieving real-world images to overcome the knowledge gaps of generative models. Specifically, to integrate missing memory for unseen novel object generation, we train a reflective retriever by self-reflective contrastive learning, which injects the generator's knowledge into the sef-reflective negatives, ensuring that the retrieved augmented images compensate for the model's missing knowledge. Furthermore, the real-object-based framework integrates finegrained visual knowledge for the generative models, tackling the distortion problem and improving the realism for fine-grained object generation. Our Real-RAG is superior in its modular application to all types of state-of-the-art text-to-image generative models and also delivers remarkable performance boosts with all of them, such as a gain of 16.18% FID score with the auto-regressive model on the Stanford Car benchmark. Figure 1. (a) The pipeline of text-to-image generative models. (b) The framework of existing retrieval-augmented methods. (c) The framework of our proposed RealRAG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 被引用 25 次
- Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language ModelsXin Zou, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu 等ICML 2025 · 被引用 1 次
- ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion ModelsPeijie Qiu, Hariharan Ramshankar, Arnau Ramisa, Amit C C 等CVPR 2026 · 被引用 1 次
- Visual Personalization Turing TestRameen Abdal, James Burgess, Sergey Tulyakov, Kuan-Chieh Jackson WangCVPR 2026
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
相关 Paper
- Reference-Based Super-Resolution via Image-Based Retrieval-Augmented Generation DiffusionByeonghun Lee, Hyunmin Cho, Hong Gyu Choi, Soo Min Kang 等ICCV 2025 · 被引用 2 次
- AR-RAG: Autoregressive Retrieval Augmentation for Image GenerationJingyuan Qi, Zhiyang Xu, Qifan Wang, Lifu HuangNeurIPS 2025 · 被引用 7 次
- Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and OptimizationJiaqi Wei, Hao Zhou, Xiang Zhang, Di Zhang 等NeurIPS 2025 · 被引用 14 次
- RAGDiffusion: Faithful Cloth Generation via External Knowledge AssimilationYuhan Li, Xianfeng Tan, Wenxiang Shang, Yubo Wu 等ICCV 2025 · 被引用 11 次
- DisarmRAG: Stealthy Retriever Poisoning to Disable Self-Correction in Retrieval-Augmented GenerationYanbo Dai, Zhenlan Ji), Zongjie Li, Kuan Li 等CCS 2026 · 被引用 2 次
