Self-supervised Learning is More Robust to Dataset Imbalance
Hong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu Ma
Abstract
Self-supervised learning (SSL) is a scalable way to learn general visual representations since it learns without labels. However, large-scale unlabeled datasets in the wild often have long-tailed label distributions, where we know little about the behavior of SSL. In this work, we systematically investigate self-supervised learning under dataset imbalance. First, we find out via extensive experiments that off-the-shelf self-supervised representations are already more robust to class imbalance than supervised representations. The performance gap between balanced and imbalanced pre-training with SSL is significantly smaller than the gap with supervised learning, across sample sizes, for both in-domain and, especially, out-of-domain evaluation. Second, towards understanding the robustness of SSL, we hypothesize that SSL learns richer features from frequent data: it may learn label-irrelevant-but-transferable features that help classify the rare classes and downstream tasks. In contrast, supervised learning has no incentive to learn features irrelevant to the labels from frequent examples. We validate this hypothesis with semi-synthetic experiments and theoretical analyses on a simplified setting. Third, inspired by the theoretical insights, we devise a re-weighted regularization technique that consistently improves the SSL representation quality on imbalanced datasets with several evaluation criteria, closing the small gap between balanced and imbalanced datasets with the same number of examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e9e3bf8-bca8-4b8f-b07e-1be63a125cc1Cited by top-tier papers66
- Improving Out-of-Distribution Robustness via Selective AugmentationHuaxiu Yao, Yu Wang, Sai Li, Linjun Zhang et al.ICML 2022 · 275 citations
- Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure AnalysisSiyi Tang, Jared Dunnmon, Khaled Kamal Saab, Xuan Zhang et al.ICLR 2022 · 157 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- Assessing the State of Self-Supervised Human Activity Recognition Using WearablesHarish Haresamudram, Irfan Essa, Thomas PlötzUbiComp 2022 · 104 citations
- Discover and Cure: Concept-aware Mitigation of Spurious CorrelationShirley Wu, Mert Yüksekgönül, Linjun Zhang, James ZouICML 2023 · 97 citations
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- Rethinking the Value of Labels for Improving Class-Imbalanced LearningYuzhe Yang, Zhi XuNeurIPS 2020 · 512 citations
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan et al.ICLR 2021 · 296 citations
- Combating Representation Learning Disparity with Geometric HarmonizationZhihan Zhou, Jiangchao Yao, Feng Hong, Ya Zhang et al.NeurIPS 2023 · 20 citations
- On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail LearningJianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao et al.ICLR 2023 · 6 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
