PreFair: Privately Generating Justifiably Fair Synthetic Data
David Pujol, Amir Gilad, Ashwin Machanavajjhala
摘要
When a database is protected by Differential Privacy (DP), its usability is limited in scope. In this scenario, generating a synthetic version of the data that mimics the properties of the private data allows users to perform any operation on the synthetic data, while maintaining the privacy of the original data. Therefore, multiple works have been devoted to devising systems for DP synthetic data generation. However, such systems may preserve or even magnify properties of the data that make it unfair, rendering the synthetic data unfit for use. In this work, we present PreFair, a system that allows for DP fair synthetic data generation. PreFair extends the state-of-the-art DP data generation mechanisms by incorporating a causal fairness criterion that ensures fair synthetic data. We adapt the notion of justifiable fairness to fit the synthetic data generation scenario. We further study the problem of generating DP fair synthetic data, showing its intractability and designing algorithms that are optimal under certain assumptions. We also provide an extensive experimental evaluation, showing that PreFair generates synthetic data that is significantly fairer than the data generated by leading DP data generation mechanisms, while remaining faithful to the private data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CuTS: Customizable Tabular Synthetic Data GenerationMark Vero, Mislav Balunovic, Martin T. VechevICML 2024 · 被引用 13 次
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential PrivacyVishnu Vinod, Krishna Pillutla, Abhradeep Guha ThakurtaNeurIPS 2025 · 被引用 12 次
- Privacy-Enhanced Database Synthesis for Benchmark PublishingYunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong 等VLDB 2025 · 被引用 3 次
- A Bayesian Nonparametric Framework for Private, Fair, and Balanced Tabular Data SynthesisForough Fazeli-Asl, Michael Minyi Zhang, Linglong Kong, Bei JiangICLR 2026
- Measuring Database Unfairness via Dependency Quantification Under Differential PrivacyMariia Vologdin, Yuchao Tao, Amir GiladVLDB 2026
它引用的顶会 Paper8
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private GeneratorsDingfan Chen, Tribhuvanesh Orekondy, Mario FritzNeurIPS 2020 · 被引用 228 次
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 被引用 174 次
- Differentially Private Query Release Through Adaptive ProjectionSergül Aydöre, William Brown, Michael Kearns, Krishnaram Kenthapadi 等ICML 2021 · 被引用 78 次
- Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic DataGeorgi Ganev, Bristena Oprisanu, Emiliano De CristofaroICML 2022 · 被引用 78 次
相关 Paper
- CausalPre: Scalable and Effective Data Pre-Processing for Causal FairnessYing Zheng, Yangfan Jiang, Kian-Lee TanICDE 2026
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao 等USENIX Security 2024 · 被引用 23 次
- DPImageBench: A Unified Benchmark for Differentially Private Image SynthesisChen Gong, Kecen Li, Zinan Lin, Tianhao WangCCS 2025 · 被引用 1 次
- Optimal Domain-Aware Privacy Mechanisms for Synthetic Data GenerationSajani Vithana, Sangwon Jung, Haoyang Hu, Viveck Cadambe 等ICML 2026
- HeteroFedSyn: Differentially Private Tabular Data Synthesis for Heterogeneous Federated SettingsXiaochen Li, Fengyu Gao, Xizixiang Wei, Tianhao Wang 等SIGMOD 2026
