RA-TTA: Retrieval-Augmented Test-Time Adaptation for Vision-Language Models
Youngjun Lee, Doyoung Kim, Junhyeok Kang, Jihwan Bang, Hwanjun Song, Jae-Gil Lee
Abstract
Vision-language models (VLMs) are known to be susceptible to distribution shifts between pre-training data and test data, and test-time adaptation (TTA) methods for VLMs have been proposed to mitigate the detrimental impact of the distribution shifts. However, the existing methods solely rely on the internal knowledge encoded within the model parameters, which are constrained to pre-training data. To complement the limitation of the internal knowledge, we propose Retrieval-Augmented-TTA (RA-TTA) for adapting VLMs to test distribution using external knowledge obtained from a web-scale image database. By fully exploiting the bi-modality of VLMs, RA-TTA adaptively retrieves proper external images for each test image to refine VLMs' predictions using the retrieved external images, where fine-grained text descriptions are leveraged to extend the granularity of external knowledge. Extensive experiments on 17 datasets demonstrate that the proposed RA-TTA outperforms the state-of-the-art methods by 3.01-9.63% on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43735de4-6358-445b-a5fe-488f6b642ce0Cited by top-tier papers4
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingZihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng et al.NeurIPS 2025 · 16 citations
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsXinqi Fan, Xueli Chen, Luoxiao Yang, Chuin Hong Yap et al.ICCV 2025 · 4 citations
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?Tilemachos Aravanis, Vladan Stojnic, Bill Psomas, Nikos Komodakis et al.CVPR 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
- PatAug: Augmentation of Augmentation for Test-Time AdaptationXinyao Li, Dan Zhang, Zhekai Du, Lei Zhu et al.ACM MM 2025 · 2 citations
- Bilateral Information-aware Test-time Adaptation for Vision-Language ModelsJingwei Sun, Jianing Zhu, Jiangchao Yao, Gang Niu et al.ICLR 2026 · 2 citations
- Efficient Test-Time Scaling for Small Vision-Language ModelsMehmet Onurcan Kaya, Desmond Elliott, Dim PapadopoulosICLR 2026 · 13 citations
- DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time AdaptationZichen Liu, Hongbo Sun, Yuxin Peng, Jiahuan ZhouAAAI 2024 · 14 citations
