Fair Densities via Boosting the Sufficient Statistics of Exponential Families
Alexander Soen, Hisham Husain, Richard Nock
Abstract
We introduce a boosting algorithm to pre-process data for fairness. Starting from an initial fair but inaccurate distribution, our approach shifts towards better data fitting while still ensuring a minimal fairness guarantee. To do so, it learns the sufficient statistics of an exponential family with boosting-compliant convergence. Importantly, we are able to theoretically prove that the learned distribution will have a representation rate and statistical rate data fairness guarantee. Unlike recent optimization based pre-processing methods, our approach can be easily adapted for continuous domain features. Furthermore, when the weak learners are specified to be decision trees, the sufficient statistics of the learned distribution can be examined to provide clues on sources of (un)fairness. Empirical results are present to display the quality of result on real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bad7bc06-4ad4-4df9-9cc2-14e9ea41ce99Builds on8
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settingsJessica Schrouff, Natalie Harris, Sanmi Koyejo, Ibrahim M. Alabdulmohsin et al.NeurIPS 2022 · 84 citations
- Rényi Fair InferenceSina Baharlouei, Maher Nouiehed, Ahmad Beirami, Meisam RazaviyaynICLR 2020 · 69 citations
- Use Privacy in Data-Driven Systems: Theory and Experiments with Machine Learnt ProgramsAnupam Datta, Matthew Fredrikson, Gihyuk Ko, Piotr Mardziel et al.CCS 2017 · 63 citations
- Data preprocessing to mitigate bias: A maximum entropy based approachL. Elisa Celis, Vijay Keswani, Nisheeth K. VishnoiICML 2020 · 45 citations
Related papers
- CausalPre: Scalable and Effective Data Pre-Processing for Causal FairnessYing Zheng, Yangfan Jiang, Kian-Lee TanICDE 2026
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
- Improving Fair Training under Correlation ShiftsYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICML 2023 · 22 citations
- Counterfactual Fairness Through Transforming Data Orthogonal to BiasShuyi Chen, Shixiang ZhuKDD 2025
- Fair Wrapping for Black-box PredictionsAlexander Soen, Ibrahim M. Alabdulmohsin, Sanmi Koyejo, Yishay Mansour et al.NeurIPS 2022 · 8 citations
