Unbiased Binning for Fairness-aware Attribute Representation
Abolfazl Asudeh, Zeinab Asoodeh, Bita Asoodeh, Omid Asudeh
Abstract
Discretizing raw features into bucketized attributes is a common step before sharing a dataset. However, this process can inadvertently introduce bias and amplify unfairness in downstream tasks.
In this paper, we address this issue by formulating the unbiased binning problem, which seeks bucketized attributes that satisfy group parity. We develop an e!cient dynamic programming algorithm to solve this problem for equal-size binning. In practice, however, an unbiased binning may incur a high price of fairness or may not exist at all, particularly when group distributions di"er substantially. To accommodate settings in which small deviations from perfect parity are acceptable, we introduce 𝐿-biased binning, which restricts group disparities across buckets to at most 𝐿. We #rst present a dynamic programming algorithm, DP, that computes the optimal solution in quadratic time. Although polynomial, DP does not scale to large datasets. To address this limitation, we propose a practically scalable algorithm based on a local search strategy (LS).
A key component of LS is a divide-and-conquer algorithm (D&C) that #nds a solution in near-linear time. We prove that D&C always returns a valid solution whenever one exists. The LS algorithm then performs a local search, using the D&C solution as an upper bound, to #nd the optimal solution. Our LS and D&C algorithms are general and are not restricted to equal-size binning.
To complement our theoretical analysis, we conduct extensive experiments on real-world and synthetic datasets. Besides con#rming the e!ciency of the algorithms, our experiments verify that while fairness-unaware binning can generate biased attribute representations, this bias can be signi#cantly reduced with a negligible price of fairness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning ModelsKi Hyun Tae, Steven Euijong WhangSIGMOD 2021 · 31 citations
- Approximation Algorithms for Fair Range ClusteringSèdjro Salomon Hotegni, Sepideh Mahabadi, Ali VakilianICML 2023 · 25 citations
- Through the Fairness Lens: Experimental Analysis and Evaluation of Entity MatchingNima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh et al.VLDB 2023 · 21 citations
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 19 citations
- Consistent Range Approximation for Fair Predictive ModelingJiongli Zhu, Sainyam Galhotra, Nazanin Sabri, Babak SalimiVLDB 2023 · 14 citations
Related papers
- Detection of Groups with Biased Representation in RankingJinyang Li, Yuval Moskovitch, H. V. JagadishICDE 2023 · 10 citations
- Fair Clustering Under a Bounded CostSeyed A. Esmaeili, Brian Brubach, Aravind Srinivasan, John DickersonNeurIPS 2021 · 36 citations
- FairHash: A Fair and Memory/Time-efficient HashmapNima Shahbazi, Stavros Sintos, Abolfazl AsudehSIGMOD 2024 · 2 citations
- Generalized Demographic Parity for Group FairnessZhimeng Jiang, Xiaotian Han, Chao Fan, Fan Yang et al.ICLR 2022 · 71 citations
- Fairness-Aware Data Preparation for Entity MatchingNima Shahbazi, Jin Wang, Zhengjie Miao, Nikita BhutaniICDE 2024 · 6 citations
