Preserving Missing Data Distribution in Synthetic Data
Xinyue Wang, Hafiz Salman Asif, Jaideep Vaidya
Abstract
Data from Web artifacts and from the Web is often sensitive and cannot be directly shared for data analysis. Therefore, synthetic data generated from the real data is increasingly used as a privacy-preserving substitute. In many cases, real data from the web has missing values where the missingness itself possesses important informational content, which domain experts leverage to improve their analysis. However, this information content is lost if either imputation or deletion is used before synthetic data generation. In this paper, we propose several methods to generate synthetic data that preserve both the observable and the missing data distributions. An extensive empirical evaluation over a range of carefully fabricated and real world datasets demonstrates the effectiveness of our approach.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Differentially Private Data Generation with Missing DataShubhankar Mohapatra, Jianqiao Zong, Florian Kerschbaum, Xi HeVLDB 2024 · 7 citations
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 56 citations
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 20 citations
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
