Debiasing Synthetic Data Generated by Deep Generative Models
Alexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Johan Decruyenaere, Christiaan Polet, Thomas Demeester, Stijn Vansteelandt
摘要
While synthetic data hold great promise for privacy protection, their statistical analysis poses significant challenges that necessitate innovative solutions. The use of deep generative models (DGMs) for synthetic data generation is known to induce considerable bias and imprecision into synthetic data analyses, compromising their inferential utility as opposed to original data analyses. This bias and uncertainty can be substantial enough to impede statistical convergence rates, even in seemingly straightforward analyses like mean calculation. The standard errors of such estimators then exhibit slower shrinkage with sample size than the typical 1 over root- rate. This complicates fundamental calculations like p-values and confidence intervals, with no straightforward remedy currently available. In response to these challenges, we propose a new strategy that targets synthetic data created by DGMs for specific data analyses. Drawing insights from debiased and targeted machine learning, our approach accounts for biases, enhances convergence rates, and facilitates the calculation of estimators with easily approximated large sample variances. We exemplify our proposal through a simulation study on toy data and two case studies on real-world data, highlighting the importance of tailoring DGMs for targeted data analysis. This debiasing strategy contributes to advancing the reliability and applicability of synthetic data in statistical inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Synthetic Data, Real Errors: How (Not) to Publish and Use Synthetic DataBoris van Breugel, Zhaozhi Qian, Mihaela van der SchaarICML 2023 · 被引用 45 次
- Empirical Gateaux Derivatives for Causal InferenceMichael I. Jordan, Yixin Wang, Angela ZhouNeurIPS 2022 · 被引用 13 次
- Automated Efficient Estimation using Monte Carlo Efficient Influence FunctionsRaj Agrawal, Sam Witty, Andy Zane, Elias BinghamNeurIPS 2024 · 被引用 6 次
相关 Paper
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 被引用 37 次
- Graphical vs. Deep Generative Models: Measuring the Impact of Differentially Private Mechanisms and Budgets on UtilityGeorgi Ganev, Kai Xu, Emiliano De CristofaroCCS 2024 · 被引用 5 次
- SoK: Privacy-Preserving Data SynthesisYuzheng Hu, Fan Wu, Qinbin Li, Yunhui Long 等S&P 2024 · 被引用 61 次
- PEARL: Data Synthesis via Private Embeddings and Adversarial Reconstruction LearningSeng Pei Liew, Tsubasa Takahashi, Michihiko UenoICLR 2022 · 被引用 32 次
- High-dimensional Analysis of Synthetic Data SelectionParham Rezaei, Filip Kovacevic, Francesco Locatello, Marco MondelliICLR 2026 · 被引用 6 次
