A study on the distribution of social biases in self-supervised learning visual models
Kirill Sirotkin, Pablo Carballeira, Marcos Escudero-Viñolo
Abstract
Deep neural networks are efficient at learning the data distribution if it is sufficiently sampled. However, they can be strongly biased by non-relevant factors implicitly incorporated in the training data. These include operational biases, such as ineffective or uneven data sampling, but also ethical concerns, as the social biases are implicitly present—even inadvertently, in the training data or explicitly defined in unfair training schedules. In tasks having impact on human processes, the learning of social biases may produce discriminatory, unethical and untrustworthy consequences. It is often assumed that social biases stem from supervised learning on labelled data, and thus, Self-Supervised Learning (SSL) wrongly appears as an efficient and bias-free solution, as it does not require labelled data. However, it was recently proven that a popular SSL method also incorporates biases. In this paper, we study the biases of a varied set of SSL visual models, trained using ImageNet data, using a method and dataset designed by psychological experts to measure social biases. We show that there is a correlation between the type of the SSL model and the number of biases that it incorporates. Furthermore, the results also suggest that this number does not strictly depend on the model's accuracy and changes throughout the network. Finally, we conclude that a careful SSL model selection process can reduce the number of social biases in the deployed model, whilst keeping high performance. The code is available at https://github.com/vpulab/SB-SSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 50 citations
- A Multidimensional Analysis of Social Biases in Vision TransformersJannik Brinkmann, Paul Swoboda, Christian BarteltICCV 2023 · 13 citations
- Target Bias Is All You Need: Zero-Shot Debiasing of Vision-Language Models With Bias CorpusTaeuk Jang, Hoin Jung, Xiaoqian WangICCV 2025 · 5 citations
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Rating Distribution Calibration for Selection Bias Mitigation in RecommendationsHaochen Liu, Da Tang, Ji Yang, Xiangyu Zhao et al.WWW 2022 · 31 citations
- Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language ModelsYi Zhou, Danushka Bollegala, José Camacho-ColladosEMNLP 2024 · 3 citations
- Self-Supervised Debiasing Using Low Rank RegularizationGeon Yeong Park, Chanyong Jung, Sangmin Lee, Jong Chul Ye et al.CVPR 2024 · 2 citations
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 190 citations
- Using Self-supervised Learning Can Improve Model FairnessSofia Yfantidou, Dimitris Spathis, Marios Constantinides, Athena Vakali et al.KDD 2024 · 3 citations
