Mitigating Subgroup Unfairness in Machine Learning Classifiers: A Data-Driven Approach
Yin Lin, Samika Gupta, H. V. Jagadish
Abstract
Fairness in machine learning, particularly in classifiers, is receiving increasing attention. However, most studies on this topic focus on fairness metrics for a limited number of predefined groups and do not address fairness across intersectional subgroups. In this paper, we investigate ways to improve subgroup fairness where subgroups are defined by the intersection of protected attributes. Specifically, our paper reveals the correlation between the representation bias of training data and model fairness. We demonstrate that biased sample collection due to historical biases and a lack of control over data collection can lead to unfairness in learned models. We introduce the concept of an “Implicit Biased Set (IBS)”, which refers to regions in the intersectional attribute space where positive and negative examples are not proportionately represented. For example, if our training data set has a disproportionate representation of black male recidivists, then criminal risk assessment tools are more likely to discriminate against black males, even if they are innocent. We propose an efficient pre-processing approach that initially identifies IBS and then employs techniques to remedy the data collection within IBS. Our evaluation shows that our method effectively mitigates various subgroup biases regardless of the downstream machine learning models used.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 78d3c213-a331-4e2c-afd5-ea14dbf736b0Cited by top-tier papers3
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
- CausalPre: Scalable and Effective Data Pre-Processing for Causal FairnessYing Zheng, Yangfan Jiang, Kian-Lee TanICDE 2026
- FairGB: A Fair Granular-Ball Generation Method for Data ClassificationQifen Yang, Yuhui Deng, Jiande Huang, Peng Zhou et al.ICML 2026
Related papers
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 16 citations
- Fair Without Leveling Down: A New Intersectional Fairness DefinitionGaurav Maheshwari, Aurélien Bellet, Pascal Denis, Mikaela KellerEMNLP 2023 · 3 citations
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 3 citations
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin et al.NeurIPS 2022 · 36 citations
- Fair Conformal Classification via Learning Representation-Based GroupsSenrong Xu, Yanke Zhou, Yuhao Tan, Zenan Li et al.ICLR 2026 · 1 citation
