Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
Ce Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi Xie
摘要
Test-time adaptation, which enables models to generalize to diverse data with unlabeled test samples, holds significant value in real-world scenarios. Recently, researchers have applied this setting to advanced pre-trained vision-language models (VLMs), developing approaches such as test-time prompt tuning to further extend their practical applicability. However, these methods typically focus solely on adapting VLMs from a single modality and fail to accumulate task-specific knowledge as more samples are processed. To address this, we introduce Dual Prototype Evolving (DPE), a novel test-time adaptation approach for VLMs that effectively accumulates task-specific knowledge from multi-modalities. Specifically, we create and evolve two sets of prototypes--textual and visual--to progressively capture more accurate multi-modal representations for target classes during test time. Moreover, to promote consistent multi-modal representations, we introduce and optimize learnable residuals for each test sample to align the prototypes from both modalities. Extensive experimental results on 15 benchmark datasets demonstrate that our proposed DPE consistently outperforms previous state-of-the-art methods while also exhibiting competitive computational efficiency. Code is available at https://github.com/zhangce01/DPE-CLIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial ReasoningZhanpeng Luo, Ce Zhang, Silong Yong, Cunxi Dai 等ICLR 2026 · 被引用 15 次
- Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated LearningWenxuan Bao, Ruxi Deng, Ruizhong Qiu, Tianxin Wei 等ICCV 2025 · 被引用 13 次
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim 等NeurIPS 2025 · 被引用 10 次
- Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common CorruptionsWenxuan Bao, Ruxi Deng, Jingrui HeNeurIPS 2025 · 被引用 7 次
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation ModelsZhanxuan Hu, Qiyu Xu, Yu Duan, Yonghang Tai 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- Dynamic Multimodal Prototype Learning in Vision-Language ModelsXingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li 等ICCV 2025
- DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time AdaptationZichen Liu, Hongbo Sun, Yuxin Peng, Jiahuan ZhouAAAI 2024 · 被引用 14 次
- Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time AdaptationQiang Zhang, Mengsheng Zhao, Jiawei Liu, Fanrui Zhang 等CVPR 2025
- SwapPrompt: Test-Time Prompt Adaptation for Vision-Language ModelsXiaosong Ma, Jie Zhang, Song Guo, Wenchao XuNeurIPS 2023 · 被引用 76 次
- PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language ModelsSi Chen, Yujia Chen, Xiaotian Yin, Xin Liu 等ACM MM 2025 · 被引用 1 次
