Vision-Language Model Selection and Reuse for Downstream Adaptation
Hao-Zhe Tan, Zhi Zhou, Yufeng Li, Lan-Zhe Guo
Abstract
Pre-trained Vision-Language Models (VLMs) are becoming increasingly popular across various visual tasks, and several open-sourced VLM variants have been released. However, selecting the best-performing pre-trained VLM for a specific downstream task is challenging since no single VLM can achieve promising performance on all downstream tasks, and evaluating all available VLMs is impossible due to time and data limitations. To address this problem, this paper proposes a novel paradigm to select and reuse VLM for downstream tasks, called Model Label Learning (MLL). The proposal contains three key modules: model labeling, which assigns labels to each VLM to describe their specialty and utility; model selection, which matches the requirements of the target task with model labels; and model reuse, which applies selected VLMs to the target task in an ensemble manner. The proposal is highly computationally efficient and growable since the model labeling process is completed target task independent and the ability could grow with the number of candidate VLMs. We also introduce a new benchmark for evaluating VLM selection methods, including 49 VLMs and 17 target task datasets. Experimental results clearly demonstrate the effectiveness of the proposed method for selecting and reusing VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7920ec5d-c789-4384-b98b-bb3f6883df53Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- VL-RouterBench: A Benchmark for Vision–Language Model RoutingZhehao Huang, Baijiong Lin, Jingyuan Zhang, Jingying Wang et al.CVPR 2026 · 6 citations
- Task Residual for Tuning Vision-Language ModelsTao Yu, Zhihe Lu, Xin Jin, Zhibo Chen et al.CVPR 2023
- Active Prompt Learning in Vision Language ModelsJihwan Bang, Sumyeong Ahn, Jae-Gil LeeCVPR 2024
- Bridge the Modality and Capability Gaps in Vision-Language Model SelectionChao Yi, Yuhang He, De-Chuan Zhan, Han-Jia YeNeurIPS 2024 · 32 citations
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha et al.CVPR 2026 · 1 citation
