On Utility and Privacy in Synthetic Genomic Data
Bristena Oprisanu, Georgi Ganev, Emiliano De Cristofaro
Abstract
The availability of genomic data is essential to progress in biomedical research, personalized medicine, etc. However, its extreme sensitivity makes it problematic, if not outright impossible, to publish or share it. As a result, several initiatives have been launched to experiment with synthetic genomic data, e.g., using generative models to learn the underlying distribution of the real data and generate artificial datasets that preserve its salient characteristics without exposing it. This paper provides the first evaluation of both utility and privacy protection of six state-of-the-art models for generating synthetic genomic data. We assess the performance of the synthetic data on several common tasks, such as allele population statistics and linkage disequilibrium. We then measure privacy through the lens of membership inference attacks, i.e., inferring whether a record was part of the training data. Our experiments show that no single approach to generate synthetic genomic data yields both high utility and strong privacy across the board. Also, the size and nature of the training dataset matter. Moreover, while some combinations of datasets and models produce synthetic data with distributions close to the real data, there often are target data points that are vulnerable to membership inference. Looking forward, our techniques can be used by practitioners to assess the risks of deploying synthetic genomic data in the wild and serve as a benchmark for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a60da07-bbfd-4263-8dc6-bf10fc385418Cited by top-tier papers3
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 37 citations
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 2 citations
- The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic DatasetsGeorgi Ganev, Emiliano De CristofaroS&P 2025
Builds on6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 543 citations
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private GeneratorsDingfan Chen, Tribhuvanesh Orekondy, Mario FritzNeurIPS 2020 · 228 citations
- Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership InferenceKlas Leino, Matt FredriksonUSENIX Security 2020
Related papers
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
- SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image GenerationYunsung Chung, Yunbei Zhang, Nassir Marrouche, Jihun HammUSENIX Security 2025
- PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic DataXinyuan Zhao, Hanlin Gu, Guibao Song, Gongxi Zhu et al.CVPR 2026
- Membership Privacy in MicroRNA-based StudiesMichael Backes, Pascal Berrang, Mathias Humbert, Praveen ManoharanCCS 2016 · 154 citations
- Inference Attacks Against Graph Generative Diffusion ModelsXiuling Wang, Xin Huang, Guibo Luo, Jianliang XuUSENIX Security 2026
