O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models
Ashshak Sharifdeen, Muhammad Akhtar Munir, Sanoojan Baliah, Salman Khan, Muhammad Haris Khan
摘要
Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although testtime prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubts on the reliability and trustworthiness of these models. Notably, more attention needs to be devoted to calibrating the test-time prompt tuning in visionlanguage models. To this end, we propose a new approach, called O-TPT that introduces orthogonality constraints on the textual features corresponding to the learnable prompts for calibrating test-time prompt tuning in VLMs. Towards introducing orthogonality constraints, we make the following contributions. First, we uncover new insights behind the suboptimal calibration performance of existing methods relying on textual feature dispersion. Second, we show that imposing a simple orthogonalization of textual features is a more effective approach towards obtaining textual dispersion. We conduct extensive experiments on various datasets with different backbones and baselines. The results indicate that our method consistently outperforms the prior state of the art in significantly reducing the overall average calibration error. Also, our method surpasses the zeroshot calibration performance on fine-grained classification tasks. Our code is available at https://github.com/ ashshaksharifdeen/O-TPT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt PretrainingHyeonseo Jang, Jaebyeong Jeon, Joong-Won Hwang, Kibok LeeCVPR 2026 · 被引用 2 次
- When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial RobustnessSunoh Kim, Daeho UmCVPR 2026 · 被引用 2 次
- SoC: Semantic Orthogonal Calibration for Test-Time Prompt TuningLeo Fillioux, Omprakash Chakraborty, Ismail Ben Ayed, Paul-Henry Cournède 等CVPR 2026 · 被引用 2 次
- InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object DetectionXiaofan Que, Dingrong Wang, Xumin Liu, Qi YuCVPR 2026
- Long-tailed Test-Time Adaptation for Vision-Language ModelsXucong Wang, Zhe Zhao, Zekun Wang, Xiaofeng Cao 等ICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
相关 Paper
- A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language ModelsShihab Aaqil Ahamed, Udaya Sampath K. Perera Miriya Thanthrige, Ranga Rodrigo, Muhammad Haris KhanICLR 2026 · 被引用 5 次
- C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature DispersionHee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Mark A. Hasegawa-Johnson 等ICLR 2024 · 被引用 84 次
- Doubly Debiased Test-Time Prompt Tuning for Vision-Language ModelsFei Song, Yi Li, Rui Wang, Jiahuan Zhou 等AAAI 2026
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
- DynaPrompt: Dynamic Test-Time Prompt TuningZehao Xiao, Shilin Yan, Jack Hong, Jiayin Cai 等ICLR 2025
