How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer Model
Hee Bin Yoo, Sungyoon Lee, Cheongjae Jang, Dong-Sig Han, Jaein Kim, Seunghyeon Lim, Byoung-Tak Zhang
Abstract
Neural networks learn effective feature representations, which can be transferred to new tasks without additional training. While larger datasets are known to improve feature transfer, the theoretical conditions for the success of such transfer remain unclear. This work investigates feature transfer in networks trained for classification to identify the conditions that enable effective clustering in unseen classes. We first reveal that higher similarity between training and unseen distributions leads to improved Cohesion and Separability. We then show that feature expressiveness is enhanced when inputs are similar to the training classes, while the features of irrelevant inputs remain indistinguishable. We validate our analysis on synthetic and benchmark datasets, including CAR, CUB, SOP, ISC, and ImageNet. Our analysis highlights the importance of the similarity between training classes and the input distribution for successful feature transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03dd46a4-9a6e-48a8-a7a4-4cb375a42883Builds on43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- On the Role of Neural Collapse in Transfer LearningTomer Galanti, András György, Marcus HutterICLR 2022 · 114 citations
- Why Do Better Loss Functions Lead to Less Transferable Features?Simon Kornblith, Ting Chen, Honglak Lee, Mohammad NorouziNeurIPS 2021 · 113 citations
- Class-relation Knowledge Distillation for Novel Class DiscoveryPeiyan Gu, Chuyu Zhang, Ruijie Xu, Xuming HeICCV 2023 · 37 citations
- Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural RepresentationsYongyi Yang, Jacob Steinhardt, Wei HuICML 2023 · 12 citations
- What Makes Transfer Learning Work for Medical Images: Feature Reuse & Other FactorsChristos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg et al.CVPR 2022 · 93 citations
