Scalable Transfer Learning with Expert Models
Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cédric Renggli, André Susano Pinto, Sylvain Gelly, Daniel Keysers, Neil Houlsby
摘要
Transfer of pre-trained representations can improve sample efficiency and reduce computational requirements for new tasks. However, representations used for transfer are usually generic, and are not tailored to a particular distribution of downstream tasks. We explore the use of expert representations for transfer with a simple, yet effective, strategy. We train a diverse set of experts by exploiting existing label structures, and use cheap-to-compute performance proxies to select the relevant expert for each target task. This strategy scales the process of transferring to new tasks, since it does not revisit the pre-training data during transfer. Accordingly, it requires little extra compute per target task, and results in a speed-up of 2-3 orders of magnitude compared to competing approaches. Further, we provide an adapter-based architecture able to compress many experts into a single model. We evaluate our approach on two different data sources and demonstrate that it outperforms baselines on over 20 diverse vision tasks in both cases. * Equal contribution. Order decided by a coin toss. † Work done while interning at Google Research. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 被引用 135 次
- Learning a Universal Template for Few-shot Dataset GeneralizationEleni Triantafillou, Hugo Larochelle, Richard S. Zemel, Vincent DumoulinICML 2021 · 被引用 113 次
- Sensitivity-Aware Visual Parameter-Efficient Fine-TuningHaoyu He, Jianfei Cai, Jing Zhang, Dacheng Tao 等ICCV 2023 · 被引用 97 次
- Ensembling Off-the-shelf Models for GAN TrainingNupur Kumari, Richard Zhang, Eli Shechtman, Jun-Yan ZhuCVPR 2022 · 被引用 75 次
它引用的顶会 Paper2
相关 Paper
- Transferring Knowledge From Large Foundation Models to Small Downstream ModelsShikai Qiu, Boran Han, Danielle C. Maddix, Shuai Zhang 等ICML 2024 · 被引用 9 次
- Efficient Adaptation of Large Vision Transformer via Adapter Re-ComposingWei Dong, Dawei Yan, Zhijun Lin, Peng WangNeurIPS 2023 · 被引用 51 次
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- Knowledge Distillation as Efficient Pre-training: Faster Convergence, Higher Data-efficiency, and Better TransferabilityRuifei He, Shuyang Sun, Jihan Yang, Song Bai 等CVPR 2022 · 被引用 40 次
- Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceTao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie 等NeurIPS 2023 · 被引用 103 次
