Retrieval-Enhanced Contrastive Vision-Text Models
Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid
Abstract
Contrastive image-text models such as CLIP form the building blocks of many state-of-the-art systems. While they excel at recognizing common generic concepts, they still struggle on fine-grained entities which are rare, or even absent from the pre-training dataset. Hence, a key ingredient to their success has been the use of large-scale curated pre-training data aiming at expanding the set of concepts that they can memorize during the pre-training stage. In this work, we explore an alternative to encoding fine-grained knowledge directly into the model's parameters: we instead train the model to retrieve this knowledge from an external memory. Specifically, we propose to equip existing vision-text models with the ability to refine their embedding with cross-modal retrieved information from a memory at inference time, which greatly improves their zero-shot predictions. Remarkably, we show that this can be done with a light-weight, single-layer, fusion transformer on top of a frozen CLIP. Our experiments validate that our retrieval-enhanced contrastive (RECO) training improves CLIP performance substantially on several challenging fine-grained tasks: for example +10.9 on Stanford Cars, +10.2 on CUB-2011 and +7.3 on the recent OVEN benchmark, where we even outperform the fine-tuned models on unseen classes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72808f43-41ce-4cef-94aa-2c16df56b355Cited by top-tier papers21
- Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image RetrievalYucheng Suo, Fan Ma, Linchao Zhu, Yi YangCVPR 2024 · 20 citations
- Retrieval-Augmented Egocentric Video CaptioningJilan Xu, Yifei Huang, Junlin Hou, Guo Chen et al.CVPR 2024 · 16 citations
- Understanding Retrieval-Augmented Task Adaptation for Vision-Language ModelsYifei Ming, Yixuan LiICML 2024 · 14 citations
- SuperCLIP: CLIP with Simple Classification SupervisionWeiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang WangNeurIPS 2025 · 6 citations
- REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing LearningJian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong et al.KDD 2025 · 5 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-TrainingChen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng et al.CVPR 2023
- Learning Customized Visual Models with Retrieval-Augmented KnowledgeHaotian Liu, Kilho Son, Jianwei Yang, Ce Liu et al.CVPR 2023
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
- MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced TrainingPavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli et al.CVPR 2024 · 29 citations
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
