A Bayesian Nonparametric Framework for Private, Fair, and Balanced Tabular Data Synthesis
Forough Fazeli-Asl, Michael Minyi Zhang, Linglong Kong, Bei Jiang
摘要
A fundamental challenge in data synthesis is protecting the fairness and privacy of the individual, particularly in data-scarce environments where underrepresented groups are at risk of further marginalization by reproducing the biases inherent in the data modeling process. We introduce a privacy-and fairness-aware generative model, which fuses the conditional generator within the framework of Bayesian nonparametric learning (BNPL). This conditional structure imposes fairness constraints in our generative model by minimizing the mutual information between generated outcomes and protected attributes. Unlike existing methods that primarily focus on sensitive binary-valued attributes, our framework extends seamlessly to non-binary attributes. Moreover, our method provides a systematic solution to class imbalance, ensuring adequate representation of underrepresented protected groups. Our proposed approach offers a scalable, privacy-preserving framework for ethical and equitable data generation, which we demonstrate by theoretical guarantees and extensive experiments on sensitive empirical examples. Randomized Response Mechanism (RRM): Randomized response is a privacy-preserving mechanism used to privatize categorical data (Wang et al., 2016). Let X be a categorical random variable taking values from a discrete set [[K]], where K is the number of categories. The RRM, denoted by M RRM (X; ϵ), perturbs the original value X according to a privacy budget ϵ, which controls the trade-off between privacy and accuracy. The ϵ-differential privacy mechanism is given as PRIVACY AND FAIRNESS PRESERVATION WITH BAYESIAN NONPARAMETRIC LEARNING Our proposed generative model uses BNPL (Fong et al., 2019) as a method of ensuring privacy and fairness protection by resampling the data from a Dirichlet process (DirP) posterior, which we will first introduce in this section. Corollary 1 (Global Perfect Privacy) Under the conditions of Proposition 1, as a → ∞, we have (i) ϵ glo → 0; moreover, (ii) δ glo p -→ 0 for fixed |W | = N -1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 被引用 174 次
- FR-Train: A Mutual Information-Based Approach to Fair and Robust TrainingYuji Roh, Kangwook Lee, Steven Whang, Changho SuhICML 2020 · 被引用 90 次
- PreFair: Privately Generating Justifiably Fair Synthetic DataDavid Pujol, Amir Gilad, Ashwin MachanavajjhalaVLDB 2023 · 被引用 16 次
相关 Paper
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 被引用 44 次
- Differentially Private and Fair Classification via Calibrated Functional MechanismJiahao Ding, Xinyue Zhang, Xiaohuan Li, Junyi Wang 等AAAI 2020 · 被引用 50 次
- Differentially Private and Fair Deep Learning: A Lagrangian Dual ApproachCuong Tran, Ferdinando Fioretto, Pascal Van HentenryckAAAI 2021 · 被引用 90 次
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 被引用 78 次
- Correct-by-Construction: Certified Individual Fairness through Neural Network TrainingRuihan Zhang, Jun SunOOPSLA 2025 · 被引用 1 次
