Why Do Better Loss Functions Lead to Less Transferable Features?
Simon Kornblith, Ting Chen, Honglak Lee, Mohammad Norouzi
摘要
Previous work has proposed many new loss functions and regularizers that improve test accuracy on image classification tasks. However, it is not clear whether these loss functions learn better representations for downstream tasks. This paper studies how the choice of training objective affects the transferability of the hidden representations of convolutional neural networks trained on ImageNet. We show that many objectives lead to statistically significant improvements in ImageNet accuracy over vanilla softmax cross-entropy, but the resulting fixed feature extractors transfer substantially worse to downstream tasks, and the choice of loss has little effect when networks are fully fine-tuned on the new tasks. Using centered kernel alignment to measure similarity between hidden representations of networks, we find that differences among loss functions are apparent only in the last few layers of the network. We delve deeper into representations of the penultimate layer, finding that different objectives and hyperparameter combinations lead to dramatically different levels of class separation. Representations with higher class separation obtain higher accuracy on the original task, but their features are less useful for downstream tasks. Our results suggest there exists a trade-off between learning invariant features for the original task and features relevant for transfer tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng 等ICML 2022 · 被引用 386 次
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You 等ICML 2022 · 被引用 122 次
- Margin-Based Few-Shot Class-Incremental Learning with Class-Level Overfitting MitigationYixiong Zou, Shanghang Zhang, Yuhua Li, Ruixuan LiNeurIPS 2022 · 被引用 100 次
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu 等NeurIPS 2022 · 被引用 93 次
- Improving neural network representations using human similarity judgmentsLukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen 等NeurIPS 2023 · 被引用 61 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
相关 Paper
- How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer ModelHee Bin Yoo, Sungyoon Lee, Cheongjae Jang, Dong-Sig Han 等NeurIPS 2025
- Does Robustness on ImageNet Transfer to Downstream Tasks?Yutaro Yamada, Mayu OtaniCVPR 2022 · 被引用 23 次
- What Do Neural Networks Learn When Trained With Random Labels?Hartmut Maennel, Ibrahim M. Alabdulmohsin, Ilya O. Tolstikhin, Robert J. N. Baldock 等NeurIPS 2020 · 被引用 99 次
- Adversarial Training Reduces Information and Improves TransferabilityMatteo Terzi, Alessandro Achille, Marco Maggipinto, Gian Antonio SustoAAAI 2021 · 被引用 25 次
- Adversarially-Trained Deep Nets Transfer Better: Illustration on Image ClassificationFrancisco Utrera, Evan Kravitz, N. Benjamin Erichson, Rajiv Khanna 等ICLR 2021 · 被引用 42 次
