Mitigating Subgroup Unfairness in Machine Learning Classifiers: A Data-Driven Approach
Yin Lin, Samika Gupta, H. V. Jagadish
摘要
Fairness in machine learning, particularly in classifiers, is receiving increasing attention. However, most studies on this topic focus on fairness metrics for a limited number of predefined groups and do not address fairness across intersectional subgroups. In this paper, we investigate ways to improve subgroup fairness where subgroups are defined by the intersection of protected attributes. Specifically, our paper reveals the correlation between the representation bias of training data and model fairness. We demonstrate that biased sample collection due to historical biases and a lack of control over data collection can lead to unfairness in learned models. We introduce the concept of an “Implicit Biased Set (IBS)”, which refers to regions in the intersectional attribute space where positive and negative examples are not proportionately represented. For example, if our training data set has a disproportionate representation of black male recidivists, then criminal risk assessment tools are more likely to discriminate against black males, even if they are innocent. We propose an efficient pre-processing approach that initially identifies IBS and then employs techniques to remedy the data collection within IBS. Our evaluation shows that our method effectively mitigates various subgroup biases regardless of the downstream machine learning models used.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
- CausalPre: Scalable and Effective Data Pre-Processing for Causal FairnessYing Zheng, Yangfan Jiang, Kian-Lee TanICDE 2026
- FairGB: A Fair Granular-Ball Generation Method for Data ClassificationQifen Yang, Yuhui Deng, Jiande Huang, Peng Zhou 等ICML 2026
相关 Paper
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 被引用 16 次
- Fair Without Leveling Down: A New Intersectional Fairness DefinitionGaurav Maheshwari, Aurélien Bellet, Pascal Denis, Mikaela KellerEMNLP 2023 · 被引用 3 次
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 被引用 3 次
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin 等NeurIPS 2022 · 被引用 36 次
- Fair Conformal Classification via Learning Representation-Based GroupsSenrong Xu, Yanke Zhou, Yuhao Tan, Zenan Li 等ICLR 2026 · 被引用 1 次
