PrivPetal: Relational Data Synthesis via Permutation Relations
Kuntai Cai, Xiaokui Xiao, Yin Yang
摘要
Releasing relational databases while preserving privacy is an important research problem with numerous applications. A canonical approach is to generate synthetic data under differential privacy (DP), which provides a strong, rigorous privacy guarantee. The problem is particularly challenging when the data involve not only entities (e.g., represented by records in tables) but also relationships (represented by foreign-key references), since if we generate random records for each entity independently, the resulting synthetic data usually fail to exhibit realistic relationships. The current state of the art, PrivLava, addresses this issue by generating random join key attributes through a sophisticated expectation-maximization (EM) algorithm. This method, however, is rather costly in terms of privacy budget consumption, due to the numerous EM iterations needed to retain high data utility. Consequently, the privacy cost of PrivLava can be prohibitive for some real-world scenarios.
We observe that the utility of the synthetic data is inherently sensitive to the join keys: changing the primary key of a record 𝑡, for example, causes 𝑡 to join with a completely different set of partner records, which may lead to a significant distribution shift of the join result. Consequently, join keys need to be kept highly accurate, meaning that enforcing DP on them inevitably incurs a high privacy cost. In this paper, we explore a different direction: synthesizing a flattened relation and subsequently decomposing it down to base relations, which eliminates the need to generate join keys. Realizing this idea is challenging, since naively flattening a relational schema leads to a rather high-dimensional table, which is hard to synthesize accurately with differential privacy.
We present a sophisticated PrivPetal approach that addresses the above issues via a novel concept: permutation relation, which is constructed as a surrogate to synthesize the flattened relation, avoiding the generation of a high-dimensional relation directly. The synthesis is done using a refined Markov random field mechanism, backed by fine-grained privacy analysis. Extensive experiments using multiple real datasets and the TPC-H benchmark demonstrate that PrivPetal significantly outperforms existing methods in terms of aggregate query accuracy on the synthetic data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- Kamino: Constraint-Aware Differentially Private Data SynthesisChang Ge, Shubhankar Mohapatra, Xi He, Ihab F. IlyasVLDB 2021 · 被引用 55 次
- PrivLava: Synthesizing Relational Data with Foreign Keys under Differential PrivacyKuntai Cai, Xiaokui Xiao, Graham CormodeSIGMOD 2023 · 被引用 25 次
- Synthesizing Linked Data Under Cardinality and Integrity ConstraintsAmir Gilad, Shweta Patwa, Ashwin MachanavajjhalaSIGMOD 2021 · 被引用 9 次
- VertiMRF: Differentially Private Vertical Federated Data SynthesisFangyuan Zhao, Zitao Li, Xuebin Ren, Bolin Ding 等KDD 2024 · 被引用 8 次
相关 Paper
- Data Synthesis via Differentially Private Markov Random FieldKuntai Cai, Xiaoyu Lei, Jianxin Wei, Xiaokui XiaoVLDB 2021 · 被引用 98 次
- Graph-Conditional Flow Matching for Relational Data GenerationDavide Scassola, Sebastiano Saccani, Luca BortolussiAAAI 2026 · 被引用 3 次
- Privacy-Enhanced Database Synthesis for Benchmark PublishingYunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong 等VLDB 2025 · 被引用 3 次
- IRG: Modular Synthetic Relational Database Generation with Complex Relational SchemasJiayu Li, Zilong Zhao, Milad Abdollahzadeh, Biplab Sikdar 等KDD 2026
- PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation ModelsVignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi 等ICML 2026 · 被引用 9 次
