Projection Head is Secretly an Information Bottleneck
Zhuo Ouyang, Kaiwen Hu, Qi Zhang, Yifei Wang, Yisen Wang
Abstract
Recently, contrastive learning has risen to be a promising paradigm for extracting meaningful data representations. Among various special designs, adding a projection head on top of the encoder during training and removing it for downstream tasks has proven to significantly enhance the performance of contrastive learning. However, despite its empirical success, the underlying mechanism of the projection head remains under-explored. In this paper, we develop an in-depth theoretical understanding of the projection head from the information-theoretic perspective. By establishing the theoretical guarantees on the downstream performance of the features before the projector, we reveal that an effective projector should act as an information bottleneck, filtering out the information irrelevant to the contrastive objective. Based on theoretical insights, we introduce modifications to projectors with training and structural regularizations. Empirically, our methods exhibit consistent improvement in the downstream performance across various real-world datasets, including CIFAR-10, CIFAR-100, and ImageNet-100. We believe our theoretical understanding on the role of the projection head will inspire more principled and advanced designs in this field. Code is available at https://github.com/PKU-ML/Projector_Theory .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2c58f6c-6342-4246-b957-6640fcc3bcb4Cited by top-tier papers8
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive LearningAchleshwar Luthra, Tianbao Yang, Tomer GalantiNeurIPS 2025 · 7 citations
- On the Alignment Between Supervised and Self-Supervised Contrastive LearningAchleshwar Luthra, Priyadarsi Mishra, Tomer GalantiICLR 2026 · 4 citations
- Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical PerspectiveYi-Ge Zhang, Jingyi Cui, Qiran Li, Yisen WangICLR 2026 · 2 citations
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
- An Augmentation-Aware Theory for Self-Supervised Contrastive LearningJingyi Cui, Hongwei Wen, Yisen WangICML 2025
Builds on26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
Related papers
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi et al.ICLR 2024 · 23 citations
- InfoNCE Induces Gaussian DistributionRoy Betser, Eyal Gofer, Meir Yossef Levi, Guy GilboaICLR 2026 · 17 citations
- Maximizing Incremental Information Entropy for Contrastive LearningJiansong Zhang, Zhuoqin Yang, Xu Wu, Xiaoling Luo et al.ICLR 2026
- The Trade-off between Universality and Label Efficiency of Representations from Contrastive LearningZhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram et al.ICLR 2023 · 1 citation
- Aligning Multimodal Representations through an Information BottleneckAntonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer et al.ICML 2025
