Alsatian: Optimizing Model Search for Deep Transfer Learning
Nils Strassenburg, Boris Glavic, Tilmann Rabl
摘要
Transfer learning is an effective technique for tuning a deep learning model when training data or computational resources are limited. Instead of training a new model from scratch, the parameters of an existing "base model" are adjusted for the new task. The accuracy of such a fine-tuned model depends on the suitability of the base model chosen. Model search automates the selection of such a base model by evaluating the suitability of candidate models for a specific task. This entails inference with each candidate model on task-specific data. With thousands of models available through model stores, the computational cost of model search is a major bottleneck for efficient transfer learning.
In this work, we present Alsatian, a novel model search system. Based on the observation that many candidate models overlap to a significant extent and following a careful bottleneck analysis, we propose optimization techniques that are applicable to many model search frameworks. These optimizations include: (i) splitting models into individual blocks that can be shared across models, (ii) caching of intermediate inference results and model blocks, and (iii) selecting a beneficial search order for models to maximize sharing of cached results. In our evaluation on state-of-the-art deep learning models from computer vision and natural language processing, we show that Alsatian outperforms baselines by up to 14×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
- LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsCuong V. Nguyen, Tal Hassner, Matthias W. Seeger, Cédric ArchambeauICML 2020 · 被引用 279 次
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 被引用 253 次
相关 Paper
- Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training DatasetsSupun Nakandala, Arun KumarSIGMOD 2022 · 被引用 6 次
- SHiFT: An Efficient, Flexible Search Engine for Transfer LearningCédric Renggli, Xiaozhe Yao, Luka Kolar, Luka Rimanic 等VLDB 2023 · 被引用 8 次
- TransTailor: Pruning the Pre-trained Model for Improved Transfer LearningBingyan Liu, Yifeng Cai, Yao Guo, Xiangqun ChenAAAI 2021 · 被引用 69 次
- Neural Data Server: A Large-Scale Search Engine for Transfer Learning DataXi Yan, David Acuna, Sanja FidlerCVPR 2020
- Which Model to Transfer? Finding the Needle in the Growing HaystackCédric Renggli, André Susano Pinto, Luka Rimanic, Joan Puigcerver 等CVPR 2022 · 被引用 13 次
