Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
Yongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie, Wenjie Li, Tat-Seng Chua
Abstract
The recent advancements in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively. Building upon this capability, we propose to enable multimodal large language models (MLLMs) to memorize and recall images within their parameters. Given a user query for visual content, the MLLM is anticipated to "recall" the relevant image from its parameters as the response. Achieving this target presents notable challenges, including inbuilt visual memory and visual recall schemes within MLLMs. To address these challenges, we introduce a generative cross-modal retrieval framework, which assigns unique identifier strings to represent images and involves two training steps: learning to memorize and learning to retrieve. The first step focuses on training the MLLM to memorize the association between images and their respective identifiers. The latter step teaches the MLLM to generate the corresponding identifier of the target image, given the textual query input. By memorizing images in MLLMs, we introduce a new paradigm to crossmodal retrieval, distinct from previous discriminative approaches. The experiments demonstrate that the generative paradigm performs effectively and efficiently even with large-scale image candidate sets. The code is released at https://github.com/liyongqi67/GRACE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aafff86-0b79-45d9-9bd8-4b850bdb3f26Cited by top-tier papers18
- Combating Multimodal LLM Hallucination via Bottom-Up Holistic ReasoningShengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang et al.AAAI 2025 · 24 citations
- When Large Vision Language Models Meet Multimodal Sequential Recommendation: An Empirical StudyPeilin Zhou, Chao Liu, Jing Ren, Xinfeng Zhou et al.WWW 2025 · 21 citations
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang et al.AAAI 2025 · 18 citations
- ContextNav: Towards Agentic Multimodal In-Context LearningHonghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang et al.ICLR 2026 · 14 citations
- Generalized Contrastive Learning for Universal Multimodal RetrievalJungsoo Lee, Janghoon Cho, Hyojin Park, Durga Malladi et al.NeurIPS 2025 · 11 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Generative Multi-Modal Knowledge Retrieval with Large Language ModelsXinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma et al.AAAI 2024 · 35 citations
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang et al.CVPR 2026 · 3 citations
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 403 citations
- TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal ModelsLeigang Qu, Haochuan Li, Tan Wang, Wenjie Wang et al.ICLR 2025
- Zero-shot Multimodal Document Retrieval via Cross-modal Question GenerationYejin Choi, Jae-Woo Park, Janghan Yoon, Saejin Kim et al.EMNLP 2025
