Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation
Sajani Vithana, Sangwon Jung, Haoyang Hu, Viveck Cadambe, Flavio Calmon, Haewon Jeong
摘要
Differential privacy (DP) imposes fundamental trade-offs between privacy and statistical fidelity in synthetic data generation. While access to public data has been shown to improve these trade-offs empirically, existing approaches use public data only indirectly, through pre-processing (e.g., using pre-trained generative models) or post-processing steps (e.g., matching target statistics estimated from public datasets), while relying on domain-agnostic DP mechanisms. In this work, we lay the theoretical framework to study the principled incorporation of public data into DP mechanisms themselves. We consider normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms. We introduce PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines. Our experiments demonstrate that PubMix significantly improves synthetic data generation quality compared to domain-agnostic privacy mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar 等ACL 2023 · 被引用 24 次
- Post-processing Private Synthetic Data for Improving Utility on Selected MeasuresHao Wang, Shivchander Sudalairaj, John Henning, Kristjan H. Greenewald 等NeurIPS 2023 · 被引用 13 次
- PrivSyn: Differentially Private Data SynthesisZhikun Zhang, Tianhao Wang, Ninghui Li, Jean Honorio 等USENIX Security 2021
- PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIsJianqing Zhang, Yang Liu, Jie Fu, Yang Hua 等ICML 2025
相关 Paper
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao 等USENIX Security 2024 · 被引用 23 次
- Epistemic Parity: Reproducibility as an Evaluation Metric for Differential PrivacyLucas Rosenblatt, Bernease Herman, Anastasia Holovenko, Wonkwon Lee 等VLDB 2023 · 被引用 11 次
- Optimal Differentially Private Model Training with Public DataAndrew Lowy, Zeman Li, Tianjian Huang, Meisam RazaviyaynICML 2024 · 被引用 9 次
- Effectively Using Public Data in Privacy Preserving Machine LearningMilad Nasr, Saeed Mahloujifar, Xinyu Tang, Prateek Mittal 等ICML 2023 · 被引用 22 次
- Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn DivergenceTianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler 等NeurIPS 2021 · 被引用 88 次
