Harnessing small projectors and multiple views for efficient vision pretraining
Arna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman, Blake A. Richards
摘要
Recent progress in self-supervised (SSL) visual representation learning has led to the development of several different proposed frameworks that rely on augmentations of images but use different loss functions. However, there are few theoretically grounded principles to guide practice, so practical implementation of each SSL framework requires several heuristics to achieve competitive performance. In this work, we build on recent analytical results to design practical recommendations for competitive and efficient SSL that are grounded in theory. Specifically, recent theory tells us that existing SSL frameworks are minimizing the same idealized loss, which is to learn features that best match the data similarity kernel defined by the augmentations used. We show how this idealized loss can be reformulated to a functionally equivalent loss that is more efficient to compute. We study the implicit bias of using gradient descent to minimize our reformulated loss function and find that using a stronger orthogonalization constraint with a reduced projector dimensionality should yield good representations. Furthermore, the theory tells us that approximating the reformulated loss should be improved by increasing the number of augmentations, and as such using multiple augmentations should lead to improved convergence. We empirically verify our findings on CIFAR, STL and Imagenet datasets, wherein we demonstrate an improved linear readout performance when training a ResNet-backbone using our theoretically grounded recommendations. Remarkably, we also demonstrate that by leveraging these insights, we can reduce the pretraining dataset size by up to 2 while maintaining downstream accuracy simply by using more data augmentations. Taken together, our work provides theoretically grounded recommendations that can be used to improve SSL convergence and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet 等ICCV 2021 · 被引用 542 次
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 被引用 467 次
相关 Paper
- Improving Self-Supervised Learning by Characterizing Idealized RepresentationsYann Dubois, Stefano Ermon, Tatsunori B. Hashimoto, Percy LiangNeurIPS 2022 · 被引用 50 次
- Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningSheng Li, Chao Wu, Ao Li, Yanzhi Wang 等ICLR 2024 · 被引用 4 次
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 等ICLR 2024 · 被引用 23 次
- Neural Harmonics: Bridging Spectral Embedding and Matrix Completion in Self-Supervised LearningMarina Munkhoeva, Ivan V. OseledetsNeurIPS 2023 · 被引用 3 次
- The SSL Interplay: Augmentations, Inductive Bias, and GeneralizationVivien Cabannes, Bobak Toussi Kiani, Randall Balestriero, Yann LeCun 等ICML 2023 · 被引用 43 次
