Data preprocessing to mitigate bias: A maximum entropy based approach
L. Elisa Celis, Vijay Keswani, Nisheeth K. Vishnoi
摘要
Data containing human or social attributes may over-or under-represent groups with respect to salient social attributes such as gender or race, which can lead to biases in downstream applications. This paper presents an algorithmic framework that can be used as a data preprocessing method towards mitigating such bias. Unlike prior work, it can efficiently learn distributions over large domains, controllably adjust the representation rates of protected groups and achieve target fairness metrics such as statistical parity, yet remains close to the empirical distribution induced by the given dataset. Our approach leverages the principle of maximum entropyamongst all distributions satisfying a given set of constraints, we should choose the one closest in KL-divergence to a given prior. While maximum entropy distributions can succinctly encode distributions over large domains, they can be difficult to compute. Our main contribution is an instantiation of this framework for our set of constraints and priors, which encode our bias mitigation goals, and that runs in time polynomial in the dimension of the data. Empirically, we observe that samples from the learned distribution have desired representation rates and statistical rates, and when used for training a classifier incurs only a slight loss in accuracy while maintaining fairness properties. * This is the full version of a paper in ICML 2020.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fair Classification with Noisy Protected Attributes: A Framework with Provable GuaranteesL. Elisa Celis, Lingxiao Huang, Vijay Keswani, Nisheeth K. VishnoiICML 2021 · 被引用 67 次
- Beyond Adult and COMPAS: Fair Multi-Class Prediction via Information ProjectionWael Alghamdi, Hsiang Hsu, Haewon Jeong, Hao Wang 等NeurIPS 2022 · 被引用 57 次
- Adaptive Sampling for Minimax Fair ClassificationShubhanshu Shekhar, Greg Fields, Mohammad Ghavamzadeh, Tara JavidiNeurIPS 2021 · 被引用 46 次
- Subgroup Robustness Grows On Trees: An Empirical Baseline InvestigationJosh Gardner, Zoran Popovic, Ludwig SchmidtNeurIPS 2022 · 被引用 27 次
- Loss Balancing for Fair Supervised LearningMohammad Mahdi Khalili, Xueru Zhang, Mahed AbroshanICML 2023 · 被引用 14 次
相关 Paper
- Fair Densities via Boosting the Sufficient Statistics of Exponential FamiliesAlexander Soen, Hisham Husain, Richard NockICML 2023 · 被引用 3 次
- Controllable Guarantees for Fair Outcomes via Contrastive Information EstimationUmang Gupta, Aaron M. Ferber, Bistra Dilkina, Greg Ver SteegAAAI 2021 · 被引用 78 次
- Controllable Universal Fair Representation LearningYue Cui, Chen Ma, Kai Zheng, Lei Chen 等WWW 2023 · 被引用 5 次
- Fair Ranking with Noisy Protected AttributesAnay Mehrotra, Nisheeth K. VishnoiNeurIPS 2022 · 被引用 24 次
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 被引用 44 次
