Investigating the Benefits of Projection Head for Representation Learning
Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi, Baharan Mirzasoleiman
Abstract
An effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2e31870-c71e-464f-bef4-15d7137aeb2dCited by top-tier papers12
- DualFed: Enjoying both Generalization and Personalization in Federated Learning via Hierachical RepresentationsGuogang Zhu, Xuefeng Liu, Jianwei Niu, Shaojie Tang et al.ACM MM 2024 · 7 citations
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive LearningAchleshwar Luthra, Tianbao Yang, Tomer GalantiNeurIPS 2025 · 7 citations
- On the Alignment Between Supervised and Self-Supervised Contrastive LearningAchleshwar Luthra, Priyadarsi Mishra, Tomer GalantiICLR 2026 · 4 citations
- CORAL: Disentangling Latent Representations in Long-Tailed DiffusionEsther Rodriguez, Monica Welfert, Samuel McDowell, Nathan Stromberg et al.NeurIPS 2025 · 1 citation
- scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell DataOlga Ovcharenko, Florian Barkmann, Philip Toma, Imant Daunhawer et al.ICML 2025
Builds on32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Projection Head is Secretly an Information BottleneckZhuo Ouyang, Kaiwen Hu, Qi Zhang, Yifei Wang et al.ICLR 2025
- Investigating Why Contrastive Learning Benefits Robustness against Label NoiseYihao Xue, Kyle Whitecross, Baharan MirzasoleimanICML 2022 · 70 citations
- Improving Self-Supervised Learning by Characterizing Idealized RepresentationsYann Dubois, Stefano Ermon, Tatsunori B. Hashimoto, Percy LiangNeurIPS 2022 · 50 citations
- Harnessing small projectors and multiple views for efficient vision pretrainingArna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman et al.NeurIPS 2024 · 5 citations
- The Mechanism of Prediction Head in Non-contrastive Self-supervised LearningZixin Wen, Yuanzhi LiNeurIPS 2022 · 44 citations
