CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation
Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers, Nicolas Thome
摘要
Vision-language models (VLMs) like CLIP exhibit strong zero-shot capabilities but often fail to generalize under distribution shifts. Test-time adaptation (TTA) allows models to update at inference time without labeled data, typically via entropy minimization. However, this objective is fundamentally misaligned with the contrastive image-text training of VLMs, limiting adaptation performance and introducing failure modes such as pseudo-label drift and class collapse. We propose CLIPTTA, a new gradient-based TTA method for vision-language models that leverages a soft contrastive loss aligned with CLIP's pre-training objective. We provide a theoretical analysis of CLIPTTA 's gradients, showing how its batchaware design mitigates the risk of collapse. We further extend CLIPTTA to the open-set setting, where both in-distribution (ID) and out-of-distribution (OOD) samples are encountered, using an Outlier Contrastive Exposure (OCE) loss to improve OOD detection. Evaluated on 75 datasets spanning diverse distribution shifts, CLIPTTA consistently outperforms entropy-based objectives and is highly competitive with state-of-the-art TTA methods, outperforming them on a large number of datasets and exhibiting more stable performance across diverse shifts. Source code is available at: CLIPTTA Repository.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image RecognitionHui Liu, Kecheng Chen, Jialiang Wang, Xianming Liu 等CVPR 2026
- STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language ModelsYiwei Fu, Hui Wan, Xiao Luo, Minghua DengCVPR 2026
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
相关 Paper
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah 等NeurIPS 2024 · 被引用 46 次
- Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot GeneralizationRan Wang, Hua Zuo, Zhen Fang, Jie LuACM MM 2024 · 被引用 7 次
- TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language ModelsJinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu 等CVPR 2026 · 被引用 2 次
- Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language ModelsShuai Zhao, Xiaohan Wang, Linchao Zhu, Yi YangICLR 2024 · 被引用 47 次
- BATCLIP: Bimodal Online Test-Time Adaptation for CLIPSarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky, Rogério Feris 等ICCV 2025 · 被引用 2 次
