To "See" is to Stereotype: Image Tagging Algorithms, Gender Recognition, and the Accuracy-Fairness Trade-off
Pinar Barlas, Kyriakos Kyriakou, Olivia Guest, Styliani Kleanthous, Jahna Otterbacher
Abstract
Machine-learned computer vision algorithms for tagging images are increasingly used by developers and researchers, having become popularized as easy-to-use "cognitive services." Yet these tools struggle with gender recognition, particularly when processing images of women, people of color and non-binary individuals. Socio-technical researchers have cited data bias as a key problem; training datasets often over-represent images of people and contexts that convey social stereotypes. The social psychology literature explains that people learn social stereotypes, in part, by observing others in particular roles and contexts, and can inadvertently learn to associate gender with scenes, occupations and activities. Thus, we study the extent to which image tagging algorithms mimic this phenomenon. We design a controlled experiment, to examine the interdependence between algorithmic recognition of context and the depicted person's gender. In the spirit of auditing to understand machine behaviors, we create a highly controlled dataset of people images, imposed on gender-stereotyped backgrounds. Our methodology is reproducible and our code publicly available. Evaluating five proprietary algorithms, we find that in three, gender inference is hindered when a background is introduced. Of the two that "see" both backgrounds and gender, it is the one whose output is most consistent with human stereotyping processes that is superior in recognizing gender. We discuss the accuracy--fairness trade-off, as well as the importance of auditing black boxes in better understanding this double-edged sword.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dea381c-2716-474a-831b-5a2d2a25f2c1Cited by top-tier papers4
- Cruising Queer HCI on the DL: A Literature Review of LGBTQ+ People in HCIJordan Taylor, Ellen Simpson, Anh-Ton Tran, Jed R. Brubaker et al.CHI 2024 · 55 citations
- RES: A Robust Framework for Guiding Visual ExplanationYuyang Gao, Tong Steven Sun, Guangji Bai, Siyi Gu et al.KDD 2022 · 29 citations
- MEDebiaser: A Human-AI Feedback System for Mitigating Bias in Multi-label Medical Image ClassificationShaohan Shi, Yuheng Shao, Haoran Jiang, Yunjie Yao et al.UIST 2025
- Rethinking Pareto Frontier: On the Optimal Trade-offs in Fair ClassificationJunyi Chai, Shenyu Lu, Xiaoqian WangICLR 2026
Related papers
- Understanding and Evaluating Racial Biases in Image CaptioningDora Zhao, Angelina Wang, Olga RussakovskyICCV 2021 · 165 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Gender Artifacts in Visual DatasetsNicole Meister, Dora Zhao, Angelina Wang, Vikram V. Ramaswamy et al.ICCV 2023 · 37 citations
- EuroGEST: Investigating gender stereotypes in multilingual language modelsJacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor et al.EMNLP 2025
- Sensemaking in User-Driven Algorithm Auditing: A Case Study on Gender Bias in an Image Captioning ModelBehnoosh Mohammadzadeh, Jules Françoise, Michèle Gouiffès, Baptiste CaramiauxCHI 2026 · 1 citation
