Views Can Be Deceiving: Improved SSL Through Feature Space Augmentation
Kimia Hamidieh, Haoran Zhang, Swami Sankaranarayanan, Marzyeh Ghassemi
Abstract
Supervised learning methods have been found to exhibit inductive biases favoring simpler features. When such features are spuriously correlated with the label, this can result in suboptimal performance on minority subgroups. Despite the growing popularity of methods which learn from unlabeled data, the extent to which these representations rely on spurious features for prediction is unclear. In this work, we explore the impact of spurious features on Self-Supervised Learning (SSL) for visual representation learning. We first empirically show that commonly used augmentations in SSL can cause undesired invariances in the image space, and illustrate this with a simple example. We further show that classical approaches in combating spurious correlations, such as dataset re-sampling during SSL, do not consistently lead to invariant representations. Motivated by these findings, we propose LateTVG to remove spurious information from these representations during pre-training, by regularizing later layers of the encoder via pruning. We find that our method produces representations which outperform the baselines on several benchmarks, without the need for group or label information during SSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 476c0eaa-73f4-46ee-8028-b4c9e898175bCited by top-tier papers4
- Mitigating Spurious Features in Contrastive Learning with Spectral RegularizationNaghmeh Ghanooni, Waleed Mustafa, Dennis Wagner, Sophie Fellenz et al.NeurIPS 2025 · 4 citations
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame ProjectionsBerken Utku Demirel, Christian HolzNeurIPS 2025 · 1 citation
- Learning Graph Invariance by Harnessing SpuriosityTianjun Yao, Yongqiang Chen, Kai Hu, Tongliang Liu et al.ICLR 2025
- On the Out-of-Distribution Generalization of Self-Supervised LearningWenwen Qiang, Jingyao Wang, Zeen Song, Jiangmeng Li et al.ICML 2025
Builds on34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Self-Supervised Debiasing Using Low Rank RegularizationGeon Yeong Park, Chanyong Jung, Sangmin Lee, Jong Chul Ye et al.CVPR 2024 · 2 citations
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 190 citations
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesWeiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng et al.CVPR 2025
- Automatic Shortcut Removal for Self-Supervised Representation LearningMatthias Minderer, Olivier Bachem, Neil Houlsby, Michael TschannenICML 2020 · 78 citations
- Improving Transferability of Representations via Augmentation-Aware Self-SupervisionHankook Lee, Kibok Lee, Kimin Lee, Honglak Lee et al.NeurIPS 2021 · 66 citations
