Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
Fei Song, Yi Li, Rui Wang, Jiahuan Zhou, Changwen Zheng, Jiangmeng Li
摘要
Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstream tasks. In this work, we analyze the underlying causes of prompt optimization bias from both the model and data perspectives. In terms of the model, the entropy minimization objective typically focuses on reducing the entropy of model predictions while overlooking their correctness. This can result in overconfident yet incorrect outputs, thereby compromising the quality of prompt optimization. On the data side, prompts affected by optimization bias can introduce misalignment between visual and textual modalities, which further aggravates the prompt optimization bias. To this end, we propose a Doubly Debiased Test-Time Prompt Tuning method, abbreviated as D 2 TPT. Specifically, we first introduce a dynamic retrieval-augmented modulation module that retrieves high-confidence knowledge from a dynamic knowledge base using the test image feature as a query, and uses the retrieved knowledge to modulate the predictions. Guided by the refined predictions, we further develop a reliability-aware prompt optimization module that incorporates a confidencebased weighted ensemble and cross-modal consistency distillation to impose regularization constraints during prompt tuning. Extensive experiments across 15 benchmark datasets involving both natural distribution shifts and cross-datasets generalization demonstrate that D 2 TPT outperforms baselines, validating its effectiveness in mitigating prompt optimization bias.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
- Diverse Data Augmentation with Diffusions for Effective Test-time Prompt TuningChun-Mei Feng, Kai Yu, Yong Liu, Salman Khan 等ICCV 2023 · 被引用 172 次
- Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationJameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak 等NeurIPS 2023 · 被引用 147 次
相关 Paper
- DynaPrompt: Dynamic Test-Time Prompt TuningZehao Xiao, Shilin Yan, Jack Hong, Jiayin Cai 等ICLR 2025
- O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language ModelsAshshak Sharifdeen, Muhammad Akhtar Munir, Sanoojan Baliah, Salman Khan 等CVPR 2025
- A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language ModelsShihab Aaqil Ahamed, Udaya Sampath K. Perera Miriya Thanthrige, Ranga Rodrigo, Muhammad Haris KhanICLR 2026 · 被引用 5 次
- Hierarchical Variational Test-Time Prompt Generation for Zero-Shot GeneralizationZhaoyang Wu, Fang Liu, Licheng Jiao, Shuo Li 等ICCV 2025 · 被引用 2 次
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 被引用 54 次
