STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language Models
Yiwei Fu, Hui Wan, Xiao Luo, Minghua Deng
摘要
This paper studies the problem of universal prompt learning at test-time for vision-language models (VLMs), aiming to enhance the adaptation of pre-trained VLMs to noisy unlabeled target data comprising both in-distribution (ID) and our-of-distribution (OOD) instances. However, existing test-time adaptation approaches often rely on unreliable pseudo-labels due to inadequate uncertainty estimation and overlook class-specific diversity in the target domain, which may result in additional adaptation bias during test time. Motivated by this, we propose a novel framework named Separability-aware Conjugate Optimization with Prototypical Retrieval (STAR) for universal test-time prompt learning of VLMs. The core of our STAR is to incorporate a separability-aware gating mechanism into conjugate optimization for reliable pseudo-learning with OOD samples. In particular, we first compute the Fisher score to quantify the separability between ID and OOD samples, which guides our soft gating mechanism for divided training. Then, we employ conjugate optimization to derive reliable pseudo-labels of unlabeled data for test-time adaptation. To further mitigate biases in OOD detection, we maintain a dynamic memory bank which stores high-confidence samples to build class-wise prototypes, which would serve as queries for prototypical retrieval to calibrate OOD detection. Extensive experiments on multiple benchmarks demonstrate that STAR consistently outperforms competing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
相关 Paper
- TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language ModelsJinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu 等CVPR 2026 · 被引用 2 次
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 被引用 57 次
- SwapPrompt: Test-Time Prompt Adaptation for Vision-Language ModelsXiaosong Ma, Jie Zhang, Song Guo, Wenchao XuNeurIPS 2023 · 被引用 76 次
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 被引用 2 次
- USE: A Unified Self-Ensembling Framework for Test-Time Prompt TuningSiru Jiang, Jian Liang, Ran He, Tieniu TanICML 2026
