Data preprocessing to mitigate bias: A maximum entropy based approach
L. Elisa Celis, Vijay Keswani, Nisheeth K. Vishnoi
Abstract
Data containing human or social attributes may over-or under-represent groups with respect to salient social attributes such as gender or race, which can lead to biases in downstream applications. This paper presents an algorithmic framework that can be used as a data preprocessing method towards mitigating such bias. Unlike prior work, it can efficiently learn distributions over large domains, controllably adjust the representation rates of protected groups and achieve target fairness metrics such as statistical parity, yet remains close to the empirical distribution induced by the given dataset. Our approach leverages the principle of maximum entropyamongst all distributions satisfying a given set of constraints, we should choose the one closest in KL-divergence to a given prior. While maximum entropy distributions can succinctly encode distributions over large domains, they can be difficult to compute. Our main contribution is an instantiation of this framework for our set of constraints and priors, which encode our bias mitigation goals, and that runs in time polynomial in the dimension of the data. Empirically, we observe that samples from the learned distribution have desired representation rates and statistical rates, and when used for training a classifier incurs only a slight loss in accuracy while maintaining fairness properties. * This is the full version of a paper in ICML 2020.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c831cb1-8357-4f35-9075-d25dba2426bcCited by top-tier papers13
- Fair Classification with Noisy Protected Attributes: A Framework with Provable GuaranteesL. Elisa Celis, Lingxiao Huang, Vijay Keswani, Nisheeth K. VishnoiICML 2021 · 67 citations
- Beyond Adult and COMPAS: Fair Multi-Class Prediction via Information ProjectionWael Alghamdi, Hsiang Hsu, Haewon Jeong, Hao Wang et al.NeurIPS 2022 · 57 citations
- Adaptive Sampling for Minimax Fair ClassificationShubhanshu Shekhar, Greg Fields, Mohammad Ghavamzadeh, Tara JavidiNeurIPS 2021 · 46 citations
- Subgroup Robustness Grows On Trees: An Empirical Baseline InvestigationJosh Gardner, Zoran Popovic, Ludwig SchmidtNeurIPS 2022 · 27 citations
- Loss Balancing for Fair Supervised LearningMohammad Mahdi Khalili, Xueru Zhang, Mahed AbroshanICML 2023 · 14 citations
Related papers
- Fair Densities via Boosting the Sufficient Statistics of Exponential FamiliesAlexander Soen, Hisham Husain, Richard NockICML 2023 · 3 citations
- Controllable Guarantees for Fair Outcomes via Contrastive Information EstimationUmang Gupta, Aaron M. Ferber, Bistra Dilkina, Greg Ver SteegAAAI 2021 · 78 citations
- Controllable Universal Fair Representation LearningYue Cui, Chen Ma, Kai Zheng, Lei Chen et al.WWW 2023 · 5 citations
- Fair Ranking with Noisy Protected AttributesAnay Mehrotra, Nisheeth K. VishnoiNeurIPS 2022 · 24 citations
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
