Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
Siddharth Joshi, Baharan Mirzasoleiman
Abstract
Self-supervised learning (SSL) learns high-quality representations from large pools of unlabeled training data. As datasets grow larger, it becomes crucial to identify the examples that contribute the most to learning such representations. This enables efficient SSL by reducing the volume of data required. Nevertheless, quantifying the value of examples for SSL has remained an open question. In this work, we address this problem for the first time, by proving that examples that contribute the most to contrastive SSL are those that have the most similar augmentations to other examples, in expectation. We provide rigorous guarantees for the generalization performance of contrastive learning on such subsets. Through extensive experiments, we show that we can safely exclude 20% of examples from CIFAR100 and 40% from STL10 and TinyImageNet, without affecting downstream task performance. In general, subsets selected by our method outperform random subsets by over 3% across these datasets. Interestingly, we also discover the subsets that contribute the most to contrastive learning are those that contribute the least to supervised learning. Code available at https://github.com/bigml-cs-ucla/sas-data-efficient-contrastive-learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7012486-02ff-49f9-aafc-a0ffaab75aedCited by top-tier papers12
- SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small ModelsYu Yang, Siddhartha Mishra, Jeffrey N. Chiang, Baharan MirzasoleimanNeurIPS 2024 · 63 citations
- Curriculum Learning With Infant Egocentric VideosSaber Sheybani, Himanshu Hansaria, Justin Wood, Linda B. Smith et al.NeurIPS 2023 · 26 citations
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi et al.ICLR 2024 · 23 citations
- D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data PruningAdyasha Maharana, Prateek Yadav, Mohit BansalICLR 2024 · 19 citations
- Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical PerspectiveYi-Ge Zhang, Jingyi Cui, Qiran Li, Yisen WangICLR 2026 · 2 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 174 citations
- An Augmentation-Aware Theory for Self-Supervised Contrastive LearningJingyi Cui, Hongwei Wen, Yisen WangICML 2025
- ArCL: Enhancing Contrastive Learning with Augmentation-Robust RepresentationsXuyang Zhao, Tianqi Du, Yisen Wang, Jun Yao et al.ICLR 2023 · 2 citations
- Understanding the Role of Equivariance in Self-supervised LearningYifei Wang, Kaiwen Hu, Sharut Gupta, Ziyu Ye et al.NeurIPS 2024 · 10 citations
- Analyzing Data-Centric Properties for Graph Contrastive LearningPuja Trivedi, Ekdeep Singh Lubana, Mark Heimann, Danai Koutra et al.NeurIPS 2022 · 13 citations
