Divide and Contrast: Self-supervised Learning from Uncurated Data
Yonglong Tian, Olivier J. Hénaff, Aäron van den Oord
Abstract
Self-supervised learning holds promise in leveraging large amounts of unlabeled data, however much of its progress has thus far been limited to highly curated pre-training data such as ImageNet. We explore the effects of contrastive learning from larger, less-curated image datasets such as YFCC, and find there is indeed a large difference in the resulting representation quality. We hypothesize that this curation gap is due to a shift in the distribution of image classes—which is more diverse and heavy-tailed—resulting in less relevant negative samples to learn from. We test this hypothesis with a new approach, Divide and Contrast (DnC), which alternates between contrastive learning and clustering-based hard negative mining. When pretrained on less curated datasets, DnC greatly improves the performance of self-supervised learning on downstream tasks, while remaining competitive with the current state-of-the-art on curated datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers42
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma et al.NeurIPS 2023 · 336 citations
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang et al.NeurIPS 2023 · 251 citations
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- Efficient Self-supervised Vision Transformers for Representation LearningChunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao et al.ICLR 2022 · 228 citations
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 213 citations
Builds on36
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Unsupervised Pre-Training of Image Features on Non-Curated DataMathilde Caron, Piotr Bojanowski, Julien Mairal, Armand JoulinICCV 2019 · 254 citations
- When Does Contrastive Visual Representation Learning Work?Elijah Cole, Xuan Yang, Kimberly Wilber, Oisin Mac Aodha et al.CVPR 2022 · 98 citations
- Neighborhood Contrastive Learning for Novel Class DiscoveryZhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo et al.CVPR 2021
- CURE: Consistency-under-Unified Semantic Regularization for Generalized Category DiscoveryYuwei Bian, Shidong Wang, Haofeng ZhangICML 2026
- Solving Inefficiency of Self-supervised Representation LearningGuangrun Wang, Keze Wang, Guangcong Wang, Philip H. S. Torr et al.ICCV 2021 · 64 citations
