Through the Data Management Lens: Experimental Analysis and Evaluation of Fair Classification
Maliha Tashfia Islam, Anna Fariha, Alexandra Meliou, Babak Salimi
Abstract
Classification, a heavily-studied data-driven machine learning task, drives an increasing number of prediction systems involving critical human decisions such as loan approval and criminal risk assessment. However, classifiers often demonstrate discriminatory behavior, especially when presented with biased data. Consequently, fairness in classification has emerged as a high-priority research area. Data management research is showing an increasing presence and interest in topics related to data and algorithmic fairness, including the topic of fair classification. The interdisciplinary efforts in fair classification, with machine learning research having the largest presence, have resulted in a large number of fairness notions and a wide range of approaches that have not been systematically evaluated and compared. In this paper, we contribute a broad analysis of 13 fair classification approaches and additional variants, over their correctness, fairness, efficiency, scalability, robustness to data errors, sensitivity to underlying ML model, data efficiency, and stability using a variety of metrics and real-world datasets. Our analysis highlights novel insights on the impact of different metrics and high-level approach characteristics on different aspects of performance. We also discuss general principles for choosing approaches suitable for different practical settings, and identify areas where data-management-centric solutions are likely to have the most impact.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Maximizing Fair Content Spread via Edge Suggestion in Social NetworksIan P. Swift, Sana Ebrahimi, Azade Nova, Abolfazl AsudehVLDB 2022 · 19 citations
- Consistent Range Approximation for Fair Predictive ModelingJiongli Zhu, Sainyam Galhotra, Nazanin Sabri, Babak SalimiVLDB 2023 · 14 citations
- F3KM: Federated, Fair, and Fast k-meansShengkun Zhu, Quanqing Xu, Jinshan Zeng, Sheng Wang et al.SIGMOD 2024 · 8 citations
- How Far Can Fairness Constraints Help Recover From Biased Data?Mohit Sharma, Amit DeshpandeICML 2024 · 7 citations
- Prerequisite-driven Fair Clustering on Heterogeneous Information NetworksJuntao Zhang, Sheng Wang, Yuan Sun, Zhiyong PengSIGMOD 2023 · 5 citations
Builds on7
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Learning Certified Individually Fair RepresentationsAnian Ruoss, Mislav Balunovic, Marc Fischer, Martin T. VechevNeurIPS 2020 · 112 citations
- Operationalizing Individual Fairness with Pairwise Fair RepresentationsPreethi Lahoti, Krishna P. Gummadi, Gerhard WeikumVLDB 2020 · 88 citations
- Rank Aggregation Algorithms for Fair ConsensusCaitlin Kuhlman, Elke A. RundensteinerVLDB 2020 · 60 citations
- Automated Feature Engineering for Algorithmic FairnessRicardo Salazar, Felix Neutatz, Ziawasch AbedjanVLDB 2021 · 42 citations
Related papers
- Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?Nico Lässig, Melanie HerschelICDE 2025
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 96 citations
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 33 citations
- Demystifying the Optimal Fair Classifier in Multi-Class ClassificationLi Zhang, Yuyuan Li, XiaoHua Feng, Jiaming Zhang et al.ICML 2026
- Faster Fair Machine via Transferring Fairness Constraints to Virtual SamplesZhou Zhai, Lei Luo, Heng Huang, Bin GuAAAI 2023
