Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
Siddharth Joshi, Baharan Mirzasoleiman
摘要
Self-supervised learning (SSL) learns high-quality representations from large pools of unlabeled training data. As datasets grow larger, it becomes crucial to identify the examples that contribute the most to learning such representations. This enables efficient SSL by reducing the volume of data required. Nevertheless, quantifying the value of examples for SSL has remained an open question. In this work, we address this problem for the first time, by proving that examples that contribute the most to contrastive SSL are those that have the most similar augmentations to other examples, in expectation. We provide rigorous guarantees for the generalization performance of contrastive learning on such subsets. Through extensive experiments, we show that we can safely exclude 20% of examples from CIFAR100 and 40% from STL10 and TinyImageNet, without affecting downstream task performance. In general, subsets selected by our method outperform random subsets by over 3% across these datasets. Interestingly, we also discover the subsets that contribute the most to contrastive learning are those that contribute the least to supervised learning. Code available at https://github.com/bigml-cs-ucla/sas-data-efficient-contrastive-learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small ModelsYu Yang, Siddhartha Mishra, Jeffrey N. Chiang, Baharan MirzasoleimanNeurIPS 2024 · 被引用 63 次
- Curriculum Learning With Infant Egocentric VideosSaber Sheybani, Himanshu Hansaria, Justin Wood, Linda B. Smith 等NeurIPS 2023 · 被引用 26 次
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 等ICLR 2024 · 被引用 23 次
- D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data PruningAdyasha Maharana, Prateek Yadav, Mohit BansalICLR 2024 · 被引用 19 次
- Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical PerspectiveYi-Ge Zhang, Jingyi Cui, Qiran Li, Yisen WangICLR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 被引用 174 次
- An Augmentation-Aware Theory for Self-Supervised Contrastive LearningJingyi Cui, Hongwei Wen, Yisen WangICML 2025
- ArCL: Enhancing Contrastive Learning with Augmentation-Robust RepresentationsXuyang Zhao, Tianqi Du, Yisen Wang, Jun Yao 等ICLR 2023 · 被引用 2 次
- Understanding the Role of Equivariance in Self-supervised LearningYifei Wang, Kaiwen Hu, Sharut Gupta, Ziyu Ye 等NeurIPS 2024 · 被引用 10 次
- Analyzing Data-Centric Properties for Graph Contrastive LearningPuja Trivedi, Ekdeep Singh Lubana, Mark Heimann, Danai Koutra 等NeurIPS 2022 · 被引用 13 次
