Head2Toe: Utilizing Intermediate Representations for Better Transfer Learning
Utku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. Mozer
摘要
Transfer-learning methods aim to improve performance in a data-scarce target domain using a model pretrained on a data-rich source domain. A cost-efficient strategy, linear probing, involves freezing the source model and training a new classification head for the target domain. This strategy is outperformed by a more costly but stateof-the-art method-fine-tuning all parameters of the source model to the target domain-possibly because fine-tuning allows the model to leverage useful information from intermediate layers which is otherwise discarded by the previously trained later layers. We explore the hypothesis that these intermediate layers might be directly exploited. We propose a method, Head-to-Toe probing (HEAD2TOE), that selects features from all layers of the source model to train a classification head for the target domain. In evaluations on the Visual Task Adaptation Benchmark (VTAB), Head2Toe matches performance obtained with fine-tuning on average while reducing training and storage cost a hundred fold or more, but critically, for out-of-distribution transfer, Head2Toe outperforms fine-tuning 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Few-shot Image Generation via Adaptation-Aware Kernel ModulationYunqing Zhao, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man CheungNeurIPS 2022 · 被引用 55 次
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar 等ICLR 2023 · 被引用 47 次
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu 等NeurIPS 2023 · 被引用 40 次
- Differentially Private Image Classification by Learning Priors from Random ProcessesXinyu Tang, Ashwinee Panda, Vikash Sehwag, Prateek MittalNeurIPS 2023 · 被引用 34 次
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- Projected GANs Converge FasterAxel Sauer, Kashyap Chitta, Jens Müller, Andreas GeigerNeurIPS 2021 · 被引用 325 次
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 被引用 204 次
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang 等ICLR 2021 · 被引用 143 次
相关 Paper
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer LearningCheng-Hao Tu, Zheda Mai, Wei-Lun ChaoCVPR 2023
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine 等ICLR 2024 · 被引用 5 次
- Attentive Multi-Layer Fusion for Vision TransformersLaure Ciernik, Marco Morik, Lukas Thede, Luca Eyring 等ICML 2026 · 被引用 3 次
- Structured Model Probing: Empowering Efficient Transfer Learning by Structured RegularizationZhi-Fan Wu, Chaojie Mao, Xue Wang, Jianwen Jiang 等CVPR 2024 · 被引用 1 次
