Estimating Structural Disparities for Face Models
Shervin Ardeshir, Cristina Segalin, Nathan Kallus
Abstract
In machine learning, disparity metrics are often defined by measuring the difference in the performance or outcome of a model, across different sub-populations (groups) of datapoints. Thus, the inputs to disparity quantification consist of a model's predictions y, the ground-truth labels for the predictions y, and group labels g for the data points. Performance of the model for each group is calculated by comparing y and y for the datapoints within a specific group, and as a result, disparity of performance across the different groups can be calculated. In many real world scenarios however, group labels (g) may not be available at scale during training and validation time, or collecting them might not be feasible or desirable as they could often be sensitive information. As a result, evaluating disparity metrics across categorical groups would not be feasible. On the other hand, in many scenarios noisy groupings may be obtainable using some form of a proxy, which would allow measuring disparity metrics across sub-populations. Here we explore performing such analysis on computer vision models trained on human faces, and on tasks such as face attribute prediction and affect estimation. Our experiments indicate that embeddings resulting from an off-the-shelf face recognition model, could meaningfully serve as a proxy for such estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c44024e1-3c20-4a07-813c-8df23e039e9dBuilds on3
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park et al.ICCV 2019 · 285 citations
Related papers
- Quantifying and Mitigating the Impact of Label Errors on Model Disparity MetricsJulius Adebayo, Melissa Hall, Bowen Yu, Bobbie ChernICLR 2023 · 2 citations
- FACET: Fairness in Computer Vision Evaluation BenchmarkLaura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval et al.ICCV 2023 · 74 citations
- Understanding and Mitigating Accuracy Disparity in RegressionJianfeng Chi, Yuan Tian, Geoffrey J. Gordon, Han ZhaoICML 2021 · 29 citations
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
- The Disparate Benefits of Deep EnsemblesKajetan Schweighofer, Adrián Arnaiz-Rodríguez, Sepp Hochreiter, Nuria OliverICML 2025
