Dynamic Multimodal Prototype Learning in Vision-Language Models
Xingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li, Yunfan Li, Junfeng Fang, Zhicai Wang, Dongsheng Wang, Hanwang Zhang
Abstract
With the increasing attention to pre-trained vision-language models (VLMs), , CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning prototypes only in the textual modality while overlooking the ambiguous semantics in class names. These ambiguities lead to textual prototypes that are insufficient to capture visual concepts, resulting in limited performance. To address this issue, we introduce ProtoMM, a training-free framework that constructs multimodal prototypes to adapt VLMs during the test time. By viewing the prototype as a discrete distribution over the textual descriptions and visual particles, ProtoMM has the ability to combine the multimodal features for comprehensive prototype learning. More importantly, the visual particles are dynamically updated as the testing stream flows. This allows our multimodal prototypes to continually learn from the data, enhancing their generalizability in unseen scenarios. In addition, we quantify the importance of the prototypes and test images by formulating their semantic distance as an optimal transport problem. Extensive experiments on 15 zero-shot benchmarks demonstrate the effectiveness of our method, achieving a 1.03% average accuracy improvement over state-of-the-art methods on ImageNet and its variant datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4ac905b-0421-41b3-8bb6-296d2da1a3baCited by top-tier papers15
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao et al.NeurIPS 2025 · 17 citations
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang et al.ICLR 2026 · 9 citations
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation ModelsZhanxuan Hu, Qiyu Xu, Yu Duan, Yonghang Tai et al.CVPR 2026 · 6 citations
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
Related papers
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 57 citations
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- Prototype-Based Test-Time Adaptation of Vision-Language ModelsZhaohong Huang, Yuxin Zhang, Wenjing Liu, Fei Chao et al.ICML 2026
- Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language ModelsYuanwei Hu, Bo Peng, Yadan Luo, zhen fang et al.ICML 2026
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
