A Bayesian Nonparametric Framework for Private, Fair, and Balanced Tabular Data Synthesis
Forough Fazeli-Asl, Michael Minyi Zhang, Linglong Kong, Bei Jiang
Abstract
A fundamental challenge in data synthesis is protecting the fairness and privacy of the individual, particularly in data-scarce environments where underrepresented groups are at risk of further marginalization by reproducing the biases inherent in the data modeling process. We introduce a privacy-and fairness-aware generative model, which fuses the conditional generator within the framework of Bayesian nonparametric learning (BNPL). This conditional structure imposes fairness constraints in our generative model by minimizing the mutual information between generated outcomes and protected attributes. Unlike existing methods that primarily focus on sensitive binary-valued attributes, our framework extends seamlessly to non-binary attributes. Moreover, our method provides a systematic solution to class imbalance, ensuring adequate representation of underrepresented protected groups. Our proposed approach offers a scalable, privacy-preserving framework for ethical and equitable data generation, which we demonstrate by theoretical guarantees and extensive experiments on sensitive empirical examples. Randomized Response Mechanism (RRM): Randomized response is a privacy-preserving mechanism used to privatize categorical data (Wang et al., 2016). Let X be a categorical random variable taking values from a discrete set [[K]], where K is the number of categories. The RRM, denoted by M RRM (X; ϵ), perturbs the original value X according to a privacy budget ϵ, which controls the trade-off between privacy and accuracy. The ϵ-differential privacy mechanism is given as PRIVACY AND FAIRNESS PRESERVATION WITH BAYESIAN NONPARAMETRIC LEARNING Our proposed generative model uses BNPL (Fong et al., 2019) as a method of ensuring privacy and fairness protection by resampling the data from a Dirichlet process (DirP) posterior, which we will first introduce in this section. Corollary 1 (Global Perfect Privacy) Under the conditions of Proposition 1, as a → ∞, we have (i) ϵ glo → 0; moreover, (ii) δ glo p -→ 0 for fixed |W | = N -1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 237e9147-f994-4843-90f0-68ff202ddb14Builds on6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 174 citations
- FR-Train: A Mutual Information-Based Approach to Fair and Robust TrainingYuji Roh, Kangwook Lee, Steven Whang, Changho SuhICML 2020 · 90 citations
- PreFair: Privately Generating Justifiably Fair Synthetic DataDavid Pujol, Amir Gilad, Ashwin MachanavajjhalaVLDB 2023 · 16 citations
Related papers
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
- Differentially Private and Fair Classification via Calibrated Functional MechanismJiahao Ding, Xinyue Zhang, Xiaohuan Li, Junyi Wang et al.AAAI 2020 · 50 citations
- Differentially Private and Fair Deep Learning: A Lagrangian Dual ApproachCuong Tran, Ferdinando Fioretto, Pascal Van HentenryckAAAI 2021 · 90 citations
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 78 citations
- Correct-by-Construction: Certified Individual Fairness through Neural Network TrainingRuihan Zhang, Jun SunOOPSLA 2025 · 1 citation
