Scalable Transfer Learning with Expert Models
Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cédric Renggli, André Susano Pinto, Sylvain Gelly, Daniel Keysers, Neil Houlsby
Abstract
Transfer of pre-trained representations can improve sample efficiency and reduce computational requirements for new tasks. However, representations used for transfer are usually generic, and are not tailored to a particular distribution of downstream tasks. We explore the use of expert representations for transfer with a simple, yet effective, strategy. We train a diverse set of experts by exploiting existing label structures, and use cheap-to-compute performance proxies to select the relevant expert for each target task. This strategy scales the process of transferring to new tasks, since it does not revisit the pre-training data during transfer. Accordingly, it requires little extra compute per target task, and results in a speed-up of 2-3 orders of magnitude compared to competing approaches. Further, we provide an adapter-based architecture able to compress many experts into a single model. We evaluate our approach on two different data sources and demonstrate that it outperforms baselines on over 20 diverse vision tasks in both cases. * Equal contribution. Order decided by a coin toss. † Work done while interning at Google Research. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bb8602d-8ffc-4abc-a2b0-100bc9512b1cCited by top-tier papers24
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 135 citations
- Learning a Universal Template for Few-shot Dataset GeneralizationEleni Triantafillou, Hugo Larochelle, Richard S. Zemel, Vincent DumoulinICML 2021 · 113 citations
- Sensitivity-Aware Visual Parameter-Efficient Fine-TuningHaoyu He, Jianfei Cai, Jing Zhang, Dacheng Tao et al.ICCV 2023 · 97 citations
- Ensembling Off-the-shelf Models for GAN TrainingNupur Kumari, Richard Zhang, Eli Shechtman, Jun-Yan ZhuCVPR 2022 · 75 citations
Builds on2
Related papers
- Transferring Knowledge From Large Foundation Models to Small Downstream ModelsShikai Qiu, Boran Han, Danielle C. Maddix, Shuai Zhang et al.ICML 2024 · 9 citations
- Efficient Adaptation of Large Vision Transformer via Adapter Re-ComposingWei Dong, Dawei Yan, Zhijun Lin, Peng WangNeurIPS 2023 · 51 citations
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin et al.AAAI 2024 · 94 citations
- Knowledge Distillation as Efficient Pre-training: Faster Convergence, Higher Data-efficiency, and Better TransferabilityRuifei He, Shuyang Sun, Jihan Yang, Song Bai et al.CVPR 2022 · 40 citations
- Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceTao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie et al.NeurIPS 2023 · 103 citations
