Implicit Modeling for Transferability Estimation of Vision Foundation Models
Yaoyan Zheng, Huiqun Wang, Nan Zhou, Di Huang
Abstract
Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitates deployment and advances the pre-training and fine-tuning paradigm. However, existing methods often struggle to accurately assess transferability for emerging pre-trained models with diverse architectures, training strategies, and task alignments. In this work, we propose Implicit Transferability Modeling (ITM), a novel framework that implicitly models each model's intrinsic transferability, coupled with a Divide-and-Conquer Variational Approximation (DVA) strategy to efficiently approximate embedding space evolution. This design enables generalization across a broader range of models and downstream tasks. Extensive experiments on a comprehensive benchmark-spanning extensive training regimes and a wider variety of model types-demonstrate that ITM consistently outperforms existing methods in terms of stability, effectiveness, and efficiency. Code is available at https://github.com/BUAAHugeGun/ITM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Foundation Model is Efficient Multimodal Multitask Model SelectorFanqing Meng, Wenqi Shao, Zhanglin Peng, Chonghe Jiang et al.NeurIPS 2023 · 26 citations
- How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation.Prabhant Singh, Sibylle Hess, Joaquin VanschorenICLR 2026 · 2 citations
- Fast and Accurate Transferability Measurement by Evaluating Intra-class Feature VarianceHuiwen Xu, U KangICCV 2023 · 12 citations
- LEAD: Exploring Logit Space Evolution for Model SelectionZixuan Hu, Xiaotong Li, Shixiang Tang, Jun Liu et al.CVPR 2024
- Understanding the Transferability of Representations via Task-RelatednessAkshay Mehra, Yunbei Zhang, Jihun HammNeurIPS 2024 · 13 citations
