Preserving Missing Data Distribution in Synthetic Data
Xinyue Wang, Hafiz Salman Asif, Jaideep Vaidya
摘要
Data from Web artifacts and from the Web is often sensitive and cannot be directly shared for data analysis. Therefore, synthetic data generated from the real data is increasingly used as a privacy-preserving substitute. In many cases, real data from the web has missing values where the missingness itself possesses important informational content, which domain experts leverage to improve their analysis. However, this information content is lost if either imputation or deletion is used before synthetic data generation. In this paper, we propose several methods to generate synthetic data that preserve both the observable and the missing data distributions. An extensive empirical evaluation over a range of carefully fabricated and real world datasets demonstrates the effectiveness of our approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Differentially Private Data Generation with Missing DataShubhankar Mohapatra, Jianqiao Zong, Florian Kerschbaum, Xi HeVLDB 2024 · 被引用 7 次
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 被引用 56 次
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 被引用 20 次
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
