Social Bias Meets Data Bias: The Impacts of Labeling and Measurement Errors on Fairness Criteria
Yiqiao Liao, Parinaz Naghizadeh
Abstract
Although many fairness criteria have been proposed to ensure that machine learning algorithms do not exhibit or amplify our existing social biases, these algorithms are trained on datasets that can themselves be statistically biased. In this paper, we investigate the robustness of existing (demographic) fairness criteria when the algorithm is trained on biased data. We consider two forms of dataset bias: errors by prior decision makers in the labeling process, and errors in the measurement of the features of disadvantaged individuals. We analytically show that some constraints (such as Demographic Parity) can remain robust when facing certain statistical biases, while others (such as Equalized Odds) are significantly violated if trained on biased data. We provide numerical experiments based on three real-world datasets (the FICO, Adult, and German credit score datasets) supporting our analytical findings. While fairness criteria are primarily chosen under normative considerations in practice, our results show that naively applying a fairness constraint can lead to not only a loss in utility for the decision maker, but more severe unfairness when data bias exists. Thus, understanding how fairness criteria react to different forms of data bias presents a critical guideline for choosing among existing fairness criteria, or for proposing new criteria, when available datasets may be biased.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 322c73e4-210b-4396-9247-5289bae68692Cited by top-tier papers5
- Fair Machine Guidance to Enhance Fair Decision Making in Biased PeopleMingzhe Yang, Hiromi Arai, Naomi Yamashita, Yukino BabaCHI 2024 · 11 citations
- Adaptive Data Debiasing through Bounded ExplorationYifan Yang, Yang Liu, Parinaz NaghizadehNeurIPS 2022 · 9 citations
- What's in a Query: Polarity-Aware Distribution-Based Fair RankingAparna Balagopalan, Kai Wang, Olawale Salaudeen, Asia Biega et al.WWW 2025 · 1 citation
- Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and FairnessDavid Jackson, Paul Groth, Hazar HarmouchVLDB 2026
- Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense DiagnosisHuijun Lian, Zekai Sun, Keqi Chen, Yingming Gao et al.ACL 2025
Builds on3
- Robust Fairness Under Covariate ShiftAshkan Rezaei, Anqi Liu, Omid Memarrast, Brian D. ZiebartAAAI 2021 · 94 citations
- How do fair decisions fare in long-term qualification?Xueru Zhang, Ruibo Tu, Yang Liu, Mingyan Liu et al.NeurIPS 2020 · 87 citations
- Decision-Making Under Selective Labels: Optimal Finite-Domain Policies and BeyondDennis WeiICML 2021 · 19 citations
Related papers
- Fairness Transferability Subject to Bounded Distribution ShiftYatong Chen, Reilly Raab, Jialu Wang, Yang LiuNeurIPS 2022 · 40 citations
- Robust Optimization for Fairness with Noisy Protected GroupsSerena Lutong Wang, Wenshuo Guo, Harikrishna Narasimhan, Andrew Cotter et al.NeurIPS 2020 · 134 citations
- Post-hoc bias scoring is optimal for fair classificationWenlong Chen, Yegor Klochkov, Yang LiuICLR 2024 · 12 citations
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin et al.NeurIPS 2022 · 36 citations
- Group Fairness by Probabilistic Modeling with Latent Fair DecisionsYooJung Choi, Meihua Dang, Guy Van den BroeckAAAI 2021 · 43 citations
