Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic Data
Georgi Ganev, Bristena Oprisanu, Emiliano De Cristofaro
摘要
Generative models trained with Differential Privacy (DP) can be used to generate synthetic data while minimizing privacy risks. We analyze the impact of DP on these models vis-à-vis underrepresented classes/subgroups of data, specifically, studying: 1) the size of classes/subgroups in the synthetic data, and 2) the accuracy of classification tasks run on them. We also evaluate the effect of various levels of imbalance and privacy budgets. Our analysis uses three state-of-the-art DP models (PrivBayes, DP-WGAN, and PATE-GAN) and shows that DP yields opposite size distributions in the generated synthetic data. It affects the gap between the majority and minority classes/subgroups; in some cases by reducing it (a "Robin Hood" effect) and, in others, by increasing it (a "Matthew" effect). Either way, this leads to (similar) disparate impacts on the accuracy of classification tasks on the synthetic data, affecting disproportionately more the underrepresented subparts of the data. Consequently, when training models on synthetic data, one might incur the risk of treating different subpopulations unevenly, leading to unreliable or unfair conclusions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- SoK: Privacy-Preserving Data SynthesisYuzheng Hu, Fan Wu, Qinbin Li, Yunhui Long 等S&P 2024 · 被引用 61 次
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 被引用 37 次
- Differential Privacy has Bounded Impact on Fairness in ClassificationPaul Mangold, Michaël Perrot, Aurélien Bellet, Marc TommasiICML 2023 · 被引用 29 次
- PreFair: Privately Generating Justifiably Fair Synthetic DataDavid Pujol, Amir Gilad, Ashwin MachanavajjhalaVLDB 2023 · 被引用 16 次
- A Learnable Discrete-Prior Fusion Autoencoder with Contrastive Learning for Tabular Data SynthesisRongchao Zhang, Yiwei Lou, Dexuan Xu, Yongzhi Cao 等AAAI 2024 · 被引用 14 次
它引用的顶会 Paper8
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative ModelsDingfan Chen, Ning Yu, Yang Zhang, Mario FritzCCS 2020 · 被引用 278 次
- Understanding Gradient Clipping in Private SGD: A Geometric PerspectiveXiangyi Chen, Zhiwei Steven Wu, Mingyi HongNeurIPS 2020 · 被引用 254 次
相关 Paper
- Graphical vs. Deep Generative Models: Measuring the Impact of Differentially Private Mechanisms and Budgets on UtilityGeorgi Ganev, Kai Xu, Emiliano De CristofaroCCS 2024 · 被引用 5 次
- Removing Disparate Impact on Model Accuracy in Differentially Private Stochastic Gradient DescentDepeng Xu, Wei Du, Xintao WuKDD 2021 · 被引用 32 次
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao 等USENIX Security 2024 · 被引用 23 次
- Differential Privacy Under Class Imbalance: Methods and Empirical InsightsLucas Rosenblatt, Yuliia Lut, Ethan Turok, Marco Avella Medina 等ICML 2025
- Differentially Private Empirical Risk Minimization under the Fairness LensCuong Tran, My H. Dinh, Ferdinando FiorettoNeurIPS 2021 · 被引用 61 次
