Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot Generalization
Ran Wang, Hua Zuo, Zhen Fang, Jie Lu
Abstract
In the field of Vision-Language Models (VLM), the Contrastive Language-Image Pretraining (CLIP) model has yielded outstanding performance on many downstream tasks through prompt tuning. By integrating image and text representations, CLIP exhibits zero-shot generalization capabilities on unseen data. However, when new categories and distribution shifts occur, the pretrained text embeddings in CLIP may not align well with unseen images, potentially leading to a decrease in CLIP's zero-shot generalization performance. To address this issue, many existing methods use test samples to update the CLIP model during testing through a process known as Test-Time Adaptation (TTA). Previous TTA techniques, such as image augmentation, can lead to overfitting given outlying samples, while methods based on teacher-student distillation can increase memory use. Further, these methods significantly increase inference time, which is a crucial factor in the testing phase. To improve robustness, mitigate overfitting, and reduce bias toward outlying samples, we propose a novel method: Self-Text Distillation with Conjugate Pseudo-labels (SCP), designed to enhance CLIP's zero-shot generalization. SCP uses gradient information from conjugate pseudo-labels to enhance the model's robustness toward distribution shifts. It also innovates by using a fixed prompt list to distil learnable prompts from within the same model, acting as a self-regulation mechanism that minimizes overfitting. Additionally, SCP is a fully test-time adaptation method that does not require retraining. It directly improves CLIP's zero-shot generalization at test time without increasing either memory overheads or inference time. In evaluations across three zero-shot generalization scenarios, SCP surpasses existing state-of-the-art methods in performance and significantly reduces inference time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d95a4642-56ce-4cf3-a74c-877ba4971a33Cited by top-tier papers6
- Learning Robust Spectral Dynamics for Temporal Domain GeneralizationEn Yu, Jie Lu, Xiaoyu Yang, Guangquan Zhang et al.NeurIPS 2025 · 22 citations
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image DetectionKuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang et al.ACM MM 2025 · 3 citations
- Is Less More? Exploring Token Condensation as Training-Free Test-Time AdaptationZixin Wang, Dong Gong, Sen Wang, Zi Huang et al.ICCV 2025 · 1 citation
- Doubly Debiased Test-Time Prompt Tuning for Vision-Language ModelsFei Song, Yi Li, Rui Wang, Jiahuan Zhou et al.AAAI 2026
- Release the Powers of Prompt Tuning: Cross-Modality Prompt TransferNingyuan Zhang, Jie Lu, Keqiuyin Li, Zhen Fang et al.ICLR 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
Related papers
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
- USE: A Unified Self-Ensembling Framework for Test-Time Prompt TuningSiru Jiang, Jian Liang, Ran He, Tieniu TanICML 2026
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2024 · 46 citations
- SwapPrompt: Test-Time Prompt Adaptation for Vision-Language ModelsXiaosong Ma, Jie Zhang, Song Guo, Wenchao XuNeurIPS 2023 · 76 citations
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
