Investigating the Benefits of Projection Head for Representation Learning
Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi, Baharan Mirzasoleiman
摘要
An effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- DualFed: Enjoying both Generalization and Personalization in Federated Learning via Hierachical RepresentationsGuogang Zhu, Xuefeng Liu, Jianwei Niu, Shaojie Tang 等ACM MM 2024 · 被引用 7 次
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive LearningAchleshwar Luthra, Tianbao Yang, Tomer GalantiNeurIPS 2025 · 被引用 7 次
- On the Alignment Between Supervised and Self-Supervised Contrastive LearningAchleshwar Luthra, Priyadarsi Mishra, Tomer GalantiICLR 2026 · 被引用 4 次
- CORAL: Disentangling Latent Representations in Long-Tailed DiffusionEsther Rodriguez, Monica Welfert, Samuel McDowell, Nathan Stromberg 等NeurIPS 2025 · 被引用 1 次
- scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell DataOlga Ovcharenko, Florian Barkmann, Philip Toma, Imant Daunhawer 等ICML 2025
它引用的顶会 Paper32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- Projection Head is Secretly an Information BottleneckZhuo Ouyang, Kaiwen Hu, Qi Zhang, Yifei Wang 等ICLR 2025
- Investigating Why Contrastive Learning Benefits Robustness against Label NoiseYihao Xue, Kyle Whitecross, Baharan MirzasoleimanICML 2022 · 被引用 70 次
- Improving Self-Supervised Learning by Characterizing Idealized RepresentationsYann Dubois, Stefano Ermon, Tatsunori B. Hashimoto, Percy LiangNeurIPS 2022 · 被引用 50 次
- Harnessing small projectors and multiple views for efficient vision pretrainingArna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman 等NeurIPS 2024 · 被引用 5 次
- The Mechanism of Prediction Head in Non-contrastive Self-supervised LearningZixin Wen, Yuanzhi LiNeurIPS 2022 · 被引用 44 次
