Head2Toe: Utilizing Intermediate Representations for Better Transfer Learning
Utku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. Mozer
Abstract
Transfer-learning methods aim to improve performance in a data-scarce target domain using a model pretrained on a data-rich source domain. A cost-efficient strategy, linear probing, involves freezing the source model and training a new classification head for the target domain. This strategy is outperformed by a more costly but stateof-the-art method-fine-tuning all parameters of the source model to the target domain-possibly because fine-tuning allows the model to leverage useful information from intermediate layers which is otherwise discarded by the previously trained later layers. We explore the hypothesis that these intermediate layers might be directly exploited. We propose a method, Head-to-Toe probing (HEAD2TOE), that selects features from all layers of the source model to train a classification head for the target domain. In evaluations on the Visual Task Adaptation Benchmark (VTAB), Head2Toe matches performance obtained with fine-tuning on average while reducing training and storage cost a hundred fold or more, but critically, for out-of-distribution transfer, Head2Toe outperforms fine-tuning 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 979eab1a-dbbe-4449-82aa-a4e418e63329Cited by top-tier papers31
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- Few-shot Image Generation via Adaptation-Aware Kernel ModulationYunqing Zhao, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man CheungNeurIPS 2022 · 55 citations
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar et al.ICLR 2023 · 47 citations
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu et al.NeurIPS 2023 · 40 citations
- Differentially Private Image Classification by Learning Priors from Random ProcessesXinyu Tang, Ashwinee Panda, Vikash Sehwag, Prateek MittalNeurIPS 2023 · 34 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- Projected GANs Converge FasterAxel Sauer, Kashyap Chitta, Jens Müller, Andreas GeigerNeurIPS 2021 · 325 citations
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 204 citations
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang et al.ICLR 2021 · 143 citations
Related papers
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer LearningCheng-Hao Tu, Zheda Mai, Wei-Lun ChaoCVPR 2023
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine et al.ICLR 2024 · 5 citations
- Attentive Multi-Layer Fusion for Vision TransformersLaure Ciernik, Marco Morik, Lukas Thede, Luca Eyring et al.ICML 2026 · 3 citations
- Structured Model Probing: Empowering Efficient Transfer Learning by Structured RegularizationZhi-Fan Wu, Chaojie Mao, Xue Wang, Jianwen Jiang et al.CVPR 2024 · 1 citation
