Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot Generalization
Ran Wang, Hua Zuo, Zhen Fang, Jie Lu
摘要
In the field of Vision-Language Models (VLM), the Contrastive Language-Image Pretraining (CLIP) model has yielded outstanding performance on many downstream tasks through prompt tuning. By integrating image and text representations, CLIP exhibits zero-shot generalization capabilities on unseen data. However, when new categories and distribution shifts occur, the pretrained text embeddings in CLIP may not align well with unseen images, potentially leading to a decrease in CLIP's zero-shot generalization performance. To address this issue, many existing methods use test samples to update the CLIP model during testing through a process known as Test-Time Adaptation (TTA). Previous TTA techniques, such as image augmentation, can lead to overfitting given outlying samples, while methods based on teacher-student distillation can increase memory use. Further, these methods significantly increase inference time, which is a crucial factor in the testing phase. To improve robustness, mitigate overfitting, and reduce bias toward outlying samples, we propose a novel method: Self-Text Distillation with Conjugate Pseudo-labels (SCP), designed to enhance CLIP's zero-shot generalization. SCP uses gradient information from conjugate pseudo-labels to enhance the model's robustness toward distribution shifts. It also innovates by using a fixed prompt list to distil learnable prompts from within the same model, acting as a self-regulation mechanism that minimizes overfitting. Additionally, SCP is a fully test-time adaptation method that does not require retraining. It directly improves CLIP's zero-shot generalization at test time without increasing either memory overheads or inference time. In evaluations across three zero-shot generalization scenarios, SCP surpasses existing state-of-the-art methods in performance and significantly reduces inference time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Robust Spectral Dynamics for Temporal Domain GeneralizationEn Yu, Jie Lu, Xiaoyu Yang, Guangquan Zhang 等NeurIPS 2025 · 被引用 22 次
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image DetectionKuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang 等ACM MM 2025 · 被引用 3 次
- Is Less More? Exploring Token Condensation as Training-Free Test-Time AdaptationZixin Wang, Dong Gong, Sen Wang, Zi Huang 等ICCV 2025 · 被引用 1 次
- Doubly Debiased Test-Time Prompt Tuning for Vision-Language ModelsFei Song, Yi Li, Rui Wang, Jiahuan Zhou 等AAAI 2026
- Release the Powers of Prompt Tuning: Cross-Modality Prompt TransferNingyuan Zhang, Jie Lu, Keqiuyin Li, Zhen Fang 等ICLR 2025
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
相关 Paper
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers 等NeurIPS 2025 · 被引用 5 次
- USE: A Unified Self-Ensembling Framework for Test-Time Prompt TuningSiru Jiang, Jian Liang, Ran He, Tieniu TanICML 2026
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah 等NeurIPS 2024 · 被引用 46 次
- SwapPrompt: Test-Time Prompt Adaptation for Vision-Language ModelsXiaosong Ma, Jie Zhang, Song Guo, Wenchao XuNeurIPS 2023 · 被引用 76 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
