Unbiased Binning for Fairness-aware Attribute Representation
Abolfazl Asudeh, Zeinab Asoodeh, Bita Asoodeh, Omid Asudeh
摘要
Discretizing raw features into bucketized attributes is a common step before sharing a dataset. However, this process can inadvertently introduce bias and amplify unfairness in downstream tasks.
In this paper, we address this issue by formulating the unbiased binning problem, which seeks bucketized attributes that satisfy group parity. We develop an e!cient dynamic programming algorithm to solve this problem for equal-size binning. In practice, however, an unbiased binning may incur a high price of fairness or may not exist at all, particularly when group distributions di"er substantially. To accommodate settings in which small deviations from perfect parity are acceptable, we introduce 𝐿-biased binning, which restricts group disparities across buckets to at most 𝐿. We #rst present a dynamic programming algorithm, DP, that computes the optimal solution in quadratic time. Although polynomial, DP does not scale to large datasets. To address this limitation, we propose a practically scalable algorithm based on a local search strategy (LS).
A key component of LS is a divide-and-conquer algorithm (D&C) that #nds a solution in near-linear time. We prove that D&C always returns a valid solution whenever one exists. The LS algorithm then performs a local search, using the D&C solution as an upper bound, to #nd the optimal solution. Our LS and D&C algorithms are general and are not restricted to equal-size binning.
To complement our theoretical analysis, we conduct extensive experiments on real-world and synthetic datasets. Besides con#rming the e!ciency of the algorithms, our experiments verify that while fairness-unaware binning can generate biased attribute representations, this bias can be signi#cantly reduced with a negligible price of fairness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning ModelsKi Hyun Tae, Steven Euijong WhangSIGMOD 2021 · 被引用 31 次
- Approximation Algorithms for Fair Range ClusteringSèdjro Salomon Hotegni, Sepideh Mahabadi, Ali VakilianICML 2023 · 被引用 25 次
- Through the Fairness Lens: Experimental Analysis and Evaluation of Entity MatchingNima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh 等VLDB 2023 · 被引用 21 次
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 被引用 19 次
- Consistent Range Approximation for Fair Predictive ModelingJiongli Zhu, Sainyam Galhotra, Nazanin Sabri, Babak SalimiVLDB 2023 · 被引用 14 次
相关 Paper
- Detection of Groups with Biased Representation in RankingJinyang Li, Yuval Moskovitch, H. V. JagadishICDE 2023 · 被引用 10 次
- Fair Clustering Under a Bounded CostSeyed A. Esmaeili, Brian Brubach, Aravind Srinivasan, John DickersonNeurIPS 2021 · 被引用 36 次
- FairHash: A Fair and Memory/Time-efficient HashmapNima Shahbazi, Stavros Sintos, Abolfazl AsudehSIGMOD 2024 · 被引用 2 次
- Generalized Demographic Parity for Group FairnessZhimeng Jiang, Xiaotian Han, Chao Fan, Fan Yang 等ICLR 2022 · 被引用 71 次
- Fairness-Aware Data Preparation for Entity MatchingNima Shahbazi, Jin Wang, Zhengjie Miao, Nikita BhutaniICDE 2024 · 被引用 6 次
