The Mechanism of Prediction Head in Non-contrastive Self-supervised Learning
Zixin Wen, Yuanzhi Li
Abstract
Recently the surprising discovery of Bootstrap Your Own Latent (BYOL) method by Grill et al. shows the negative term in contrastive loss can be removed if we add the so-called prediction head to the network architecture, which breaks the symmetry between the positive pairs. This initiated the research of non-contrastive self-supervised learning. It is mysterious why even when trivial collapsed global optimal solutions exist, neural networks trained by (stochastic) gradient descent can still learn competitive representations and avoid collapsed solutions. This phenomenon is one of the most typical examples of implicit bias in deep learning optimization, and its underlying mechanism remains little understood to this day. In this work, we present our empirical and theoretical discoveries about the mechanism of prediction head in non-contrastive self-supervised learning methods. Empirically, we find that when the prediction head is initialized as an identity matrix with only its off-diagonal entries being trained, the network can learn competitive representations even though the trivial optima still exist in the training objective. Moreover, we observe a consistent rise and fall trajectory of off-diagonal entries during training. Our evidence suggests that understanding the identity-initialized prediction head is a good starting point for understanding the mechanism of the trainable prediction head. Theoretically, we present a framework to understand the behavior of the trainable, but identity-initialized prediction head. Under a simple setting, we characterized the substitution effect and acceleration effect of the prediction head during the training process. The substitution effect happens when learning the stronger features in some neurons can substitute for learning these features in other neurons through updating the prediction head. And the acceleration effect happens when the substituted features can accelerate the learning of other weaker features to prevent them from being ignored. These two effects together enable the neural networks to learn all the features rather than focus only on learning the stronger features, which is likely the cause of the dimensional collapse phenomenon. To the best of our knowledge, this is also the first end-to-end optimization guarantee for non-contrastive methods using nonlinear neural networks with a trainable prediction head and normalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9efddb30-8e52-47dc-a403-3590ee34b6fcCited by top-tier papers19
- BYOL-Explore: Exploration by Bootstrapped PredictionZhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires et al.NeurIPS 2022 · 104 citations
- On the Power of Foundation ModelsYang YuanICML 2023 · 55 citations
- Understanding Self-Predictive Learning for Reinforcement LearningYunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires et al.ICML 2023 · 46 citations
- Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self-Supervised LearningHugues Van Assel, Mark Ibrahim, Tommaso Biancalani, Aviv Regev et al.NeurIPS 2025 · 39 citations
- Matrix Information Theory for Self-Supervised LearningYifan Zhang, Zhiquan Tan, Jingqin Yang, Weiran Huang et al.ICML 2024 · 26 citations
Builds on43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Implicit Contrastive Representation Learning with Guided Stop-gradientByeongchan Lee, Sehyun LeeNeurIPS 2023 · 3 citations
- Understanding self-supervised learning dynamics without contrastive pairsYuandong Tian, Xinlei Chen, Surya GanguliICML 2021 · 338 citations
- The Edge of Orthogonality: A Simple View of What Makes BYOL TickPierre Harvey Richemond, Allison C. Tam, Yunhao Tang, Florian Strub et al.ICML 2023 · 17 citations
- Towards a Unified Theoretical Understanding of Non-contrastive Learning via Rank Differential MechanismZhijian Zhuo, Yifei Wang, Jinwen Ma, Yisen WangICLR 2023 · 1 citation
- Implicit variance regularization in non-contrastive SSLManu Srinath Halvagal, Axel Laborieux, Friedemann ZenkeNeurIPS 2023 · 18 citations
