Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models
Zhihe Lu, Jiawang Bai, Xin Li, Zeyu Xiao, Xinchao Wang
摘要
Fine-tuning pre-trained vision-language models (VLMs), e.g., CLIP, for the open-world generalization has gained increasing popularity due to its practical value. However, performance advancements are limited when relying solely on intricate algorithmic designs for a single model, even one exhibiting strong performance, e.g., CLIP-ViT-B/16. This paper, for the first time, explores the collaborative potential of leveraging much weaker VLMs to enhance the generalization of a robust single model. The affirmative findings motivate us to address the generalization problem from a novel perspective, i.e., ensemble of pre-trained VLMs. We introduce three customized ensemble strategies, each tailored to one specific scenario. Firstly, we introduce the zero-shot ensemble, automatically adjusting the logits of different models based on their confidence when only pre-trained VLMs are available. Furthermore, for scenarios with extra few-shot samples, we propose the training-free and tuning ensemble, offering flexibility based on the availability of computing resources. The proposed ensemble strategies are evaluated on zero-shot, base-to-new, and cross-dataset generalization, achieving new state-of-the-art performance. Notably, this work represents an initial stride toward enhancing the generalization performance of VLMs via ensemble. The code is available at https://github.com/zhiheLu/Ensemble_VLM.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingTaolin Zhang, Jinpeng Wang, Hang Guo, Tao Dai 等NeurIPS 2024 · 被引用 30 次
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor 等EMNLP 2025 · 被引用 9 次
- TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent CollaborationYiwei Guo, Shaobin Zhuang, Kunchang Li, Yu Qiao 等NeurIPS 2024 · 被引用 9 次
- Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li 等CVPR 2026 · 被引用 7 次
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free ApproachBaoshun Tong, Hanjiang Lai, Yan Pan, Jian YinCVPR 2025
- LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and GeneralizationJingzheng Li, Hailong SunACM MM 2023 · 被引用 5 次
- Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsYongjin Yang, Jongwoo Ko, Se-Young YunEMNLP 2024 · 被引用 1 次
- CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentHaoyu Song, Li Dong, Weinan Zhang, Ting Liu 等ACL 2022
- Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias CorrectingXingyu Zhu, Beier Zhu, Yi Tan, Shuo Wang 等NeurIPS 2024 · 被引用 36 次
