Foundation Model is Efficient Multimodal Multitask Model Selector
Fanqing Meng, Wenqi Shao, Zhanglin Peng, Chonghe Jiang, Kaipeng Zhang, Yu Qiao, Ping Luo
Abstract
This paper investigates an under-explored but important problem: given a collection of pre-trained neural networks, predicting their performance on each multi-modal task without fine-tuning them, such as image recognition, referring, captioning, visual question answering, and text question answering. A brute-force approach is to finetune all models on all target datasets, bringing high computational costs. Although recent-advanced approaches employed lightweight metrics to measure models' transferability,they often depend heavily on the prior knowledge of a single task, making them inapplicable in a multi-modal multi-task scenario. To tackle this issue, we propose an efficient multi-task model selector (EMMS), which employs large-scale foundation models to transform diverse label formats such as categories, texts, and bounding boxes of different downstream tasks into a unified noisy label embedding. EMMS can estimate a model's transferability through a simple weighted linear regression, which can be efficiently solved by an alternating minimization algorithm with a convergence guarantee. Extensive experiments on 5 downstream tasks with 24 datasets show that EMMS is fast, effective, and generic enough to assess the transferability of pre-trained models, making it the first model selection method in the multi-task scenario. For instance, compared with the state-of-the-art method LogME enhanced by our label embeddings, EMMS achieves 9.0%, 26.3%, 20.1%, 54.8%, 12.2% performance gain on image recognition, referring, captioning, visual question answering, and text question answering, while bringing 5.13x, 6.29x, 3.59x, 6.19x, and 5.66x speedup in wall-clock time, respectively. The code is available at https://github.com/OpenGVLab/Multitask-Model-Selector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55a8a741-1f56-45f4-8cdd-56387a348a41Cited by top-tier papers8
- Towards Robust Multi-Modal Reasoning via Model SelectionXiangyan Liu, Rongxue Li, Wei Ji, Tao LinICLR 2024 · 9 citations
- Understanding Museum Exhibits using Vision-Language ReasoningAda-Astrid Balauca, Sanjana Garai, Stefan Balauca, Rasesh Udayakumar Shetty et al.ICCV 2025 · 3 citations
- Toward Enhancing Representation Learning in Federated Multi-Task SettingsMehdi Setayesh, Mahdi Beitollahi, Yasser H. Khalil, Hongliang LiICLR 2026 · 2 citations
- How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation.Prabhant Singh, Sibylle Hess, Joaquin VanschorenICLR 2026 · 2 citations
- Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text RankingJun Bai, Zhuofan Chen, Zhenzi Li, Hanhua Hong et al.EMNLP 2024 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 253 citations
- Vision-Language Model Selection and Reuse for Downstream AdaptationHao-Zhe Tan, Zhi Zhou, Yufeng Li, Lan-Zhe GuoICML 2025
- Implicit Modeling for Transferability Estimation of Vision Foundation ModelsYaoyan Zheng, Huiqun Wang, Nan Zhou, Di HuangNeurIPS 2025 · 1 citation
- Ranked from Within: Ranking Large Multimodal Models Without LabelsWeijie Tu, Weijian Deng, Dylan Campbell, Yu Yao et al.ICML 2025
- EMR-Merging: Tuning-Free High-Performance Model MergingChenyu Huang, Peng Ye, Tao Chen, Tong He et al.NeurIPS 2024 · 134 citations
