The hidden uniform cluster prior in self-supervised learning
Mido Assran, Randall Balestriero, Quentin Duval, Florian Bordes, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael G. Rabbat, Nicolas Ballas
摘要
A successful paradigm in representation learning is to perform self-supervised pretraining using tasks based on mini-batch statistics (e.g., SimCLR, VICReg, SwAV, MSN). We show that in the formulation of all these methods is an overlooked prior to learn features that enable uniform clustering of the data. While this prior has led to remarkably semantic representations when pretraining on class-balanced data, such as ImageNet, we demonstrate that it can hamper performance when pretraining on class-imbalanced data. By moving away from conventional uniformity priors and instead preferring power-law distributed feature clusters, we show that one can improve the quality of the learned representations on real-world class-imbalanced datasets. To demonstrate this, we develop an extension of the Masked Siamese Networks (MSN) method to support the use of arbitrary features priors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon 等ICCV 2023 · 被引用 181 次
- On the Stepwise Nature of Self-Supervised LearningJames B. Simon, Maksis Knutins, Liu Ziyin, Daniel Geisz 等ICML 2023 · 被引用 45 次
- Group Robust Classification Without Any Group InformationChristos Tsirigotis, João Monteiro, Pau Rodríguez, David Vázquez 等NeurIPS 2023 · 被引用 34 次
- Contrastive Tuning: A Little Help to Make Masked Autoencoders ForgetJohannes Lehner, Benedikt Alkin, Andreas Fürst, Elisabeth Rumetshofer 等AAAI 2024 · 被引用 28 次
- UniCorn: A Unified Contrastive Learning Approach for Multi-view Molecular Representation LearningShikun Feng, Yuyan Ni, Minghao Li, Yanwen Huang 等ICML 2024 · 被引用 22 次
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 被引用 190 次
- Improve Representation for Imbalanced Regression through Geometric ConstraintsZijian Dong, Yilei Wu, Chongyao Chen, Yingtian Zou 等CVPR 2025
- Mitigating Spurious Features in Contrastive Learning with Spectral RegularizationNaghmeh Ghanooni, Waleed Mustafa, Dennis Wagner, Sophie Fellenz 等NeurIPS 2025 · 被引用 4 次
- Train a One-Million-Way Instance Classifier for Unsupervised Visual Representation LearningYu Liu, Lianghua Huang, Pan Pan, Bin Wang 等AAAI 2021 · 被引用 3 次
- Unbiased Subclass Regularization for Semi-Supervised Semantic SegmentationDayan Guan, Jiaxing Huang, Aoran Xiao, Shijian LuCVPR 2022 · 被引用 57 次
