Privacy-Enhanced Database Synthesis for Benchmark Publishing
Yunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong, Bo Tang, Yu-Xuan Qiu, Rui Mao, Ye Yuan, Makoto Onizuka, Chuan Xiao
摘要
Benchmarking is crucial for evaluating a DBMS, yet existing benchmarks often fail to reflect the varied nature of user workloads. As a result, there is increasing momentum toward creating databases that incorporate real-world user data to more accurately mirror business environments. However, privacy concerns deter users from directly sharing their data, underscoring the importance of creating synthesized databases for benchmarking that also prioritize privacy protection. Differential privacy (DP)-based data synthesis has become a key method for safeguarding privacy when sharing data, but the focus has largely been on minimizing errors in aggregate queries or downstream ML tasks, with less attention given to benchmarking factors like query runtime performance. This paper delves into differentially private database synthesis specifically for benchmark publishing scenarios, aiming to produce a synthetic database whose benchmarking factors closely resemble those of the original data. Introducing PrivBench , an innovative synthesis framework based on sum-product networks (SPNs), we support the synthesis of high-quality benchmark databases that maintain fidelity in both data distribution and query runtime performance while preserving privacy. We validate that PrivBench can ensure database-level DP even when generating multi-relation databases with complex reference relationships. Our extensive experiments show that PrivBench efficiently synthesizes data that maintains privacy and excels in both data distribution similarity and query runtime similarity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private GeneratorsDingfan Chen, Tribhuvanesh Orekondy, Mario FritzNeurIPS 2020 · 被引用 228 次
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu 等VLDB 2022 · 被引用 169 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- Dealer: An End-to-End Model Marketplace with Differential PrivacyJinfei Liu, Jian Lou, Junxu Liu, Li Xiong 等VLDB 2021 · 被引用 99 次
相关 Paper
- DPImageBench: A Unified Benchmark for Differentially Private Image SynthesisChen Gong, Kecen Li, Zinan Lin, Tianhao WangCCS 2025 · 被引用 1 次
- PrivLava: Synthesizing Relational Data with Foreign Keys under Differential PrivacyKuntai Cai, Xiaokui Xiao, Graham CormodeSIGMOD 2023 · 被引用 25 次
- Differentially Private Sum-Product NetworksXenia Heilmann, Mattia Cerrato, Ernst AlthausICML 2024 · 被引用 1 次
- Data Synthesis via Differentially Private Markov Random FieldKuntai Cai, Xiaoyu Lei, Jianxin Wei, Xiaokui XiaoVLDB 2021 · 被引用 98 次
- Benchmarking Differentially Private Tabular Data Synthesis: [Experiments & Analysis]Kai Chen, Xiaochen Li, Chen Gong, Ryan McKenna 等SIGMOD 2026 · 被引用 4 次
