Realistic Test-Time Adaptation of Vision-Language Models
Maxime Zanella, Clément Fuchs, Christophe De Vleeschouwer, Ismail Ben Ayed
Abstract
The zero-shot capabilities of Vision-Language Models (VLMs) have been widely leveraged to improve predictive performance. However, previous works on transductive or test-time adaptation (TTA) often make strong assumptions about the data distribution, such as the presence of all classes. Our work challenges these favorable deployment scenarios and introduces a more realistic evaluation framework, including (i) a variable number of effective classes for adaptation within a single batch, and (ii) non-i.i.d. batches of test samples in online adaptation settings. We provide comprehensive evaluations, comparisons, and ablation studies that demonstrate how current transductive or TTA methods for VLMs systematically compromise the models' initial zero-shot robustness across various realistic scenarios, favoring performance gains under advantageous assumptions about the test sample distributions. Furthermore, we introduce StatA, a versatile method that can handle a wide range of deployment scenarios, including those with a variable number of effective classes at test time. Our approach incorporates a novel regularization term designed specifically for VLMs, which acts as a statistical anchor preserving the initial text-encoder knowledge, particularly in low-data regimes. Code available at https://github.com/MaxZanella/StatA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7e8fa33-be09-4845-b771-773fd61d107fCited by top-tier papers14
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation ModelsZhanxuan Hu, Qiyu Xu, Yu Duan, Yonghang Tai et al.CVPR 2026 · 6 citations
- Statistics Caching Test-Time Adaptation for Vision-Language ModelsZenghao Guan, Yucan Zhou, Wu Liu, Xiaoyan GuNeurIPS 2025 · 5 citations
- Endowing Vision-Language Models with System 2 Thinking for Fine-grained Visual RecognitionYutong Yang, Lifu Huang, Yijie Lin, Xi Peng et al.AAAI 2026 · 2 citations
- Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time TransductionJiazhen Huang, Zhiming Liu, Changhu Wang, Wei Ju et al.ICML 2026 · 2 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
- NOTE: Robust Continual Test-time Adaptation Against Temporal CorrelationTaesik Gong, Jongheon Jeong, Taewon Kim, Yewon Kim et al.NeurIPS 2022 · 227 citations
- Contrastive Test-Time AdaptationDian Chen, Dequan Wang, Trevor Darrell, Sayna EbrahimiCVPR 2022 · 219 citations
Related papers
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsYiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou et al.ACM MM 2025 · 1 citation
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
- SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute PromptingChenyu Zhang, Kunlun Xu, Zichen Liu, Yuxin Peng et al.CVPR 2025
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2024 · 46 citations
