Distributed Synthesis of Differentially Private Tabular Datasets
Yucheng Fu, Tianyao Gu, Elaine Shi, Tianhao Wang
摘要
Differentially private synthetic data generation has emerged as a powerful tool for sharing data while protecting individuals' privacy. However, when the attributes of sensitive data are distributed across multiple entities such as hospitals, companies, or government agencies, accurately generating synthetic data becomes challenging. In particular, it is difficult to capture informative statistical correlations and use them to guide data synthesis without gathering the entire private dataset. In response to this challenge, we propose a secure multi-party computation protocol for differentially private tabular data synthesis in the distributed setting. Our protocol contains two new primitives. The first is a protocol that exploits distributed point functions to efficiently estimate two-way marginals (pairwise joint distributions of attributes) across vertically distributed data. The second is a protocol for generating noise via batched lookups in the cumulative distribution function table. As a concrete demonstration, we build a distributed version of AIM, a state-of-the-art DP data-synthesis algorithm. Our implementation achieves the same utility as its centralized version while reducing end-to-end runtime by orders of magnitude compared with prior work. For example, we can synthesize the "Adult" dataset in 24 minutes in a real-world WAN setting, whereas the existing protocol is estimated to take 57 days.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- The Discrete Gaussian for Differential PrivacyClément L. Canonne, Gautam Kamath, Thomas SteinkeNeurIPS 2020 · 被引用 355 次
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure AggregationPeter Kairouz, Ziyu Liu, Thomas SteinkeICML 2021 · 被引用 291 次
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- Lightweight Techniques for Private Heavy HittersDan Boneh, Elette Boyle, Henry Corrigan-Gibbs, Niv Gilboa 等S&P 2021 · 被引用 134 次
- New Oracle-Efficient Algorithms for Private Synthetic Data ReleaseGiuseppe Vietri, Grace Tian, Mark Bun, Thomas Steinke 等ICML 2020 · 被引用 86 次
相关 Paper
- CaPS: Collaborative and Private Synthetic Data Generation from Distributed SourcesSikha Pentyala, Mayana Pereira, Martine De CockICML 2024 · 被引用 6 次
- FLAIM: AIM-based Synthetic Data Generation in the Federated SettingSamuel Maddock, Graham Cormode, Carsten MapleKDD 2024 · 被引用 5 次
- Privacy-Preserving Data Release Leveraging Optimal Transport and Particle Gradient DescentKonstantin Donhauser, Javier Abad Martinez, Neha Hulkund, Fanny YangICML 2024 · 被引用 6 次
- HeteroFedSyn: Differentially Private Tabular Data Synthesis for Heterogeneous Federated SettingsXiaochen Li, Fengyu Gao, Xizixiang Wei, Tianhao Wang 等SIGMOD 2026
- PrivSyn: Differentially Private Data SynthesisZhikun Zhang, Tianhao Wang, Ninghui Li, Jean Honorio 等USENIX Security 2021
