Self-supervised Learning is More Robust to Dataset Imbalance
Hong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu Ma
摘要
Self-supervised learning (SSL) is a scalable way to learn general visual representations since it learns without labels. However, large-scale unlabeled datasets in the wild often have long-tailed label distributions, where we know little about the behavior of SSL. In this work, we systematically investigate self-supervised learning under dataset imbalance. First, we find out via extensive experiments that off-the-shelf self-supervised representations are already more robust to class imbalance than supervised representations. The performance gap between balanced and imbalanced pre-training with SSL is significantly smaller than the gap with supervised learning, across sample sizes, for both in-domain and, especially, out-of-domain evaluation. Second, towards understanding the robustness of SSL, we hypothesize that SSL learns richer features from frequent data: it may learn label-irrelevant-but-transferable features that help classify the rare classes and downstream tasks. In contrast, supervised learning has no incentive to learn features irrelevant to the labels from frequent examples. We validate this hypothesis with semi-synthetic experiments and theoretical analyses on a simplified setting. Third, inspired by the theoretical insights, we devise a re-weighted regularization technique that consistently improves the SSL representation quality on imbalanced datasets with several evaluation criteria, closing the small gap between balanced and imbalanced datasets with the same number of examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- Improving Out-of-Distribution Robustness via Selective AugmentationHuaxiu Yao, Yu Wang, Sai Li, Linjun Zhang 等ICML 2022 · 被引用 275 次
- Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure AnalysisSiyi Tang, Jared Dunnmon, Khaled Kamal Saab, Xuan Zhang 等ICLR 2022 · 被引用 157 次
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 被引用 119 次
- Assessing the State of Self-Supervised Human Activity Recognition Using WearablesHarish Haresamudram, Irfan Essa, Thomas PlötzUbiComp 2022 · 被引用 104 次
- Discover and Cure: Concept-aware Mitigation of Spurious CorrelationShirley Wu, Mert Yüksekgönül, Linjun Zhang, James ZouICML 2023 · 被引用 97 次
它引用的顶会 Paper29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
相关 Paper
- Rethinking the Value of Labels for Improving Class-Imbalanced LearningYuzhe Yang, Zhi XuNeurIPS 2020 · 被引用 512 次
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan 等ICLR 2021 · 被引用 296 次
- Combating Representation Learning Disparity with Geometric HarmonizationZhihan Zhou, Jiangchao Yao, Feng Hong, Ya Zhang 等NeurIPS 2023 · 被引用 20 次
- On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail LearningJianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao 等ICLR 2023 · 被引用 6 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
