PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language Models
Si Chen, Yujia Chen, Xiaotian Yin, Xin Liu, Huakai Lai, Tianzhu Zhang
摘要
Leveraging Vision-Language Models (VLMs) like CLIP for various downstream tasks has emerged as a significant research trend. Recently, researchers have introduced Test-Time Adaptation (TTA) as a technique for models to learn online from unlabeled samples at test time, improving the generalization performance of VLMs to target domains. However, existing TTA methods either require expensive backpropagating gradient computations for each test sample or only extract knowledge from a limited number of historical test samples in the cache model, resulting in suboptimal adaptation performance. To address these limitations, we propose a Prototype Adaptive Fusion (PAF) framework, a novel TTA approach that makes full use of historical knowledge from test samples. Unlike traditional cache-based methods, which store only a few low-entropy samples per class, PAF introduces a prototype fusion mechanism that constructs class prototype representations through cumulatively merging features from qualified test samples. Furthermore, we propose an enhanced version, Easy-Hard PAF (EH-PAF), which adaptively applies a category-specific strategy based on CLIP prediction to improve performance. Extensive experiments across 15 diverse datasets demonstrate that our method consistently outperforms previous state-of-the-art approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language ModelsLiangsheng Liu, Si Chen, Jiamin Wu, Weiwei Feng 等ICLR 2026 · 被引用 4 次
- Beyond Logits: Coherent Hallucination Mitigation via Attention Contrastive DecodingYujia Chen, Rui Sun, Huayu Mai, Wangkai Li 等ICML 2026
相关 Paper
- Prototype-Based Test-Time Adaptation of Vision-Language ModelsZhaohong Huang, Yuxin Zhang, Wenjing Liu, Fei Chao 等ICML 2026
- Statistics Caching Test-Time Adaptation for Vision-Language ModelsZenghao Guan, Yucan Zhou, Wu Liu, Xiaoyan GuNeurIPS 2025 · 被引用 5 次
- BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingTaolin Zhang, Jinpeng Wang, Hang Guo, Tao Dai 等NeurIPS 2024 · 被引用 30 次
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah 等NeurIPS 2024 · 被引用 46 次
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers 等NeurIPS 2025 · 被引用 5 次
