Understanding the Learning Phases in Self-Supervised Learning via Critical Periods
Janghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. Lunga
Abstract
Self-supervised learning (SSL) has emerged as a powerful pretraining strategy to learn transferable representations from unlabeled data. Yet, it remains unclear how long SSL models should be pretrained to yield such representations. Contrary to the prevailing heuristic that longer pretraining translates to better downstream performance, we observe a transferability trade-off: across diverse SSL settings, intermediate checkpoints can yield stronger out-of-domain (OOD) generalization, whereas additional pretraining primarily benefits in-domain (ID) performance. From this observation, we hypothesize that SSL progresses through learning phases that can be characterized via the lens of critical periods (CP). Prior work on CP has shown that supervised models exhibit an early phase of high plasticity, followed by a consolidation phase where adaptability declines but task-specific performance increases. Since traditional CP analysis was developed for supervised settings, we rethink it for SSL in two ways. First, we inject deficits to perturb the pretraining data and assess their lasting impact on representation quality via downstream tasks. Second, we compute the Fisher Information on pretext objectives to track plasticity, quantifying how sensitive model parameters are to the pretext task. Our experiments suggest that SSL models may exhibit their own CP, with CP closure coinciding with a sweet spot for broad downstream transferability. Leveraging these insights, we introduce CP-guided checkpoint selection as a strategy for selecting checkpoints that offer stronger OOD transferability. Finally, to balance the transferability trade-off, we present CP-guided self-distillation, which selectively distills layer representations from the intermediate checkpoint into their overspecialized counterparts in the final checkpoint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f81f23d-7e57-4c76-84af-ca94c45625d9Builds on47
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
Related papers
- Critical Learning Periods for Multisensory Integration in Deep NetworksMichael Kleinman, Alessandro Achille, Stefano SoattoCVPR 2023
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 190 citations
- Self-Supervised Dataset Distillation for Transfer LearningDong Bok Lee, Seanie Lee, Joonho Ko, Kenji Kawaguchi et al.ICLR 2024 · 9 citations
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
