Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
Yifei Ming, Yixuan Li
Abstract
Pre-trained contrastive vision-language models have demonstrated remarkable performance across a wide range of tasks. However, they often struggle on fine-trained datasets with categories not adequately represented during pre-training, which makes adaptation necessary. Recent works have shown promising results by utilizing samples from web-scale databases for retrieval-augmented adaptation, especially in low-data regimes. Despite the empirical success, understanding how retrieval impacts the adaptation of vision-language models remains an open research question. In this work, we adopt a reflective perspective by presenting a systematic study to understand the roles of key components in retrieval-augmented adaptation. We unveil new insights on uni-modal and cross-modal retrieval and highlight the critical role of logit ensemble for effective adaptation. We further present theoretical underpinnings that directly support our empirical observations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4eb83946-4eac-44ee-8a62-11bc32878f86Cited by top-tier papers10
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- On the Comparison between Multi-modal and Single-modal Contrastive LearningWei Huang, Andi Han, Yongqiang Chen, Yuan Cao et al.NeurIPS 2024 · 26 citations
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsXinqi Fan, Xueli Chen, Luoxiao Yang, Chuin Hong Yap et al.ICCV 2025 · 4 citations
- Is the Modality Gap a Bug or a Feature? A Robustness PerspectiveRhea Chowers, Oshri Naparstek, Udi Barzelay, Yair WeissCVPR 2026 · 4 citations
- Reevaluating the Intra-Modal Misalignment Hypothesis in CLIPJonas Herzog, Yue WangCVPR 2026 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- RA-TTA: Retrieval-Augmented Test-Time Adaptation for Vision-Language ModelsYoungjun Lee, Doyoung Kim, Junhyeok Kang, Jihwan Bang et al.ICLR 2025
- ProLoG: Hybrid Prompt and LoRA Based Adaptation of Vision-Language Models for OOD GeneralizationJungwuk Park, Dong-Jun Han, Jaekyun MoonAAAI 2026
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
- VladVA: Discriminative Fine-tuning of LVLMsYassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis et al.CVPR 2025
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 2 citations
