Which Model to Transfer? Finding the Needle in the Growing Haystack
Cédric Renggli, André Susano Pinto, Luka Rimanic, Joan Puigcerver, Carlos Riquelme, Ce Zhang, Mario Lucic
Abstract
Transfer learning has been recently popularized as a data-efficient alternative to training models from scratch, in particular for computer vision tasks where it provides a remarkably solid baseline. The emergence of rich model repositories, such as TensorFlow Hub, enables the practitioners and researchers to unleash the potential of these models across a wide range of downstream tasks. As these repositories keep growing exponentially, efficiently selecting a good model for the task at hand becomes paramount. We provide a formalization of this problem through afamiliar notion of regret and introduce the predominant strategies, namely task-agnostic (e.g. ranking models by their ImageNet performance) and task-aware search strategies (such as linear or kNN evaluation). We conduct a large-scale empirical study and show that both task-agnostic and task-aware methods can yield high regret. We then propose a simple and computationally efficient hybrid search strategy which outperforms the existing approaches. We highlight the practical benefits of the proposed solution on a set of 19 diverse vision tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8048a807-dabb-47d9-afad-09b2e324c8b3Cited by top-tier papers10
- Model Spider: Learning to Rank Pre-Trained Models EfficientlyYi-Kai Zhang, Ting-Ji Huang, Yao-Xiang Ding, De-Chuan Zhan et al.NeurIPS 2023 · 57 citations
- Great Models Think Alike: Improving Model Reliability via Inter-Model Latent AgreementAilin Deng, Miao Xiong, Bryan HooiICML 2023 · 9 citations
- SHiFT: An Efficient, Flexible Search Engine for Transfer LearningCédric Renggli, Xiaozhe Yao, Luka Kolar, Luka Rimanic et al.VLDB 2023 · 8 citations
- What to Pre-Train on? Efficient Intermediate Task SelectionClifton Poth, Jonas Pfeiffer, Andreas Rücklé, Iryna GurevychEMNLP 2021 · 8 citations
- Automatic Feasibility Study via Data Quality Analysis for ML: A Case-Study on Label NoiseCédric Renggli, Luka Rimanic, Luka Kolar, Wentao Wu et al.ICDE 2023 · 8 citations
Builds on5
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
- Transferability and Hardness of Supervised Classification TasksAnh Tuan Tran, Cuong V. Nguyen, Tal HassnerICCV 2019 · 201 citations
- Scalable Transfer Learning with Expert ModelsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cédric Renggli et al.ICLR 2021 · 70 citations
- DEPARA: Deep Attribution Graph for Deep Knowledge TransferabilityJie Song, Yixin Chen, Jingwen Ye, Xinchao Wang et al.CVPR 2020
Related papers
- Sommelier: Curating DNN Models for the MassesPeizhen Guo, Bo Hu, Wenjun HuSIGMOD 2022 · 13 citations
- Neural Data Server: A Large-Scale Search Engine for Transfer Learning DataXi Yan, David Acuna, Sanja FidlerCVPR 2020
- Guided Recommendation for Model Fine-TuningHao Li, Charless C. Fowlkes, Hao Yang, Onkar Dabeer et al.CVPR 2023
- Hub-Pathway: Transfer Learning from A Hub of Pre-trained ModelsYang Shu, Zhangjie Cao, Ziyang Zhang, Jianmin Wang et al.NeurIPS 2022 · 7 citations
- AutoTransfer: AutoML with Knowledge Transfer - An Application to Graph Neural NetworksKaidi Cao, Jiaxuan You, Jiaju Liu, Jure LeskovecICLR 2023
