Toward Understanding the Feature Learning Process of Self-supervised Contrastive Learning
Zixin Wen, Yuanzhi Li
摘要
How can neural networks trained by contrastive learning extract features from the unlabeled data? Why does contrastive learning usually need much stronger data augmentations than supervised learning to ensure good representations? These questions involve both the optimization and statistical aspects of deep learning, but can hardly be answered by the analysis of supervised learning, where the target functions are the highest pursuit. Indeed, in self-supervised learning, it is inevitable to relate to the optimization/generalization of neural networks to how they can encode the latent structures in the data, which we refer to as the feature learning process. In this work, we formally study how contrastive learning learns the feature representations for neural networks by analyzing its feature learning process. We consider the case where our data are comprised of two types of features: the more semantically aligned sparse features which we want to learn from, and the other dense features we want to avoid. Theoretically, we prove that contrastive learning using ReLU networks provably learns the desired sparse features if proper augmentations are adopted. We present an underlying principle called feature decoupling to explain the effects of augmentations, where we theoretically characterize how augmentations can reduce the correlations of dense features between positive samples while keeping the correlations of sparse features intact, thereby forcing the neural networks to learn from the self-supervision of sparse features. Empirically, we verified that the feature decoupling principle matches the underlying mechanism of contrastive learning in practice. V1 appeared on June 1, 2021. V2 polished writing and added citations, V3 corrected related works. We would like to thank Zeyuan Allen-Zhu for many helpful suggestions on the experiments, and thank Qi Lei, Jason D. Lee for clarifying results of their paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu 等NeurIPS 2022 · 被引用 199 次
- Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding MethodsRandall Balestriero, Yann LeCunNeurIPS 2022 · 被引用 189 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 被引用 115 次
- On the Convergence and Sample Complexity Analysis of Deep Q-Networks with ε-Greedy ExplorationShuai Zhang, Hongkang Li, Meng Wang, Miao Liu 等NeurIPS 2023 · 被引用 57 次
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 被引用 425 次
相关 Paper
- An Augmentation-Aware Theory for Self-Supervised Contrastive LearningJingyi Cui, Hongwei Wen, Yisen WangICML 2025
- Learning to Imagine: Diversify Memory for Incremental Learning using Unlabeled DataYu-Ming Tang, Yi-Xing Peng, Wei-Shi ZhengCVPR 2022 · 被引用 33 次
- MetAug: Contrastive Learning via Meta Feature AugmentationJiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 等ICML 2022 · 被引用 32 次
- Self-Supervised Debiasing Using Low Rank RegularizationGeon Yeong Park, Chanyong Jung, Sangmin Lee, Jong Chul Ye 等CVPR 2024 · 被引用 2 次
- Disentangled Contrastive Learning on GraphsHaoyang Li, Xin Wang, Ziwei Zhang, Zehuan Yuan 等NeurIPS 2021 · 被引用 136 次
