USENIX Security2024Top-tier venue
A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic Data
Meenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc Rocher
Abstract
Recent advances in synthetic data generation (SDG) have been hailed as a solution to the difficult problem of sharing sensitive data while protecting privacy. SDG aims to learn statistical properties of real data in order to generate"artificial"data that are structurally and statistically similar to sensitive data. However, prior research suggests that inference attacks on synthetic data can undermine privacy, but only for specific outlier records. In this work, we introduce a new attribute inference attack against synthetic data. The attack is based on linear reconstruction methods for aggregate statistics, which target all records in the dataset, not only outliers. We evaluate our attack on state-of-the-art SDG algorithms, including Probabilistic Graphical Models, Generative Adversarial Networks, and recent differentially private SDG mechanisms. By defining a formal privacy game, we show that our attack can be highly accurate even on arbitrary records, and that this is the result of individual information leakage (as opposed to population-level inference). We then systematically evaluate the tradeoff between protecting privacy and preserving statistical utility. Our findings suggest that current SDG methods cannot consistently provide sufficient privacy protection against inference attacks while retaining reasonable utility. The best method evaluated, a differentially private SDG mechanism, can provide both protection against inference attacks and reasonable utility, but only in very specific settings. Lastly, we show that releasing a larger number of synthetic records can improve utility but at the cost of making attacks far more effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd43fcc3-dbeb-4a96-a65f-8cff1456ca18Cited by top-tier papers8
- "What do you want from theory alone?" Experimenting with Tight Auditing of Differentially Private Synthetic Data GenerationMeenatchi Sundaram Muthu Selva Annamalai, Georgi Ganev, Emiliano De CristofaroUSENIX Security 2024 · 24 citations
- Access Denied: Meaningful Data Access for Quantitative Algorithm AuditsJuliette Zaccour, Reuben Binns, Luc RocherCHI 2025 · 9 citations
- Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic DataShenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren et al.EMNLP 2025 · 7 citations
- SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority OversamplingGeorgi Ganev, MohammadReza Nazari, Rees Davison, Amirhassan Fallah Dizche et al.ICLR 2026 · 6 citations
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 2 citations
Builds on20
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated LearningMilad Nasr, Reza Shokri, Amir HoumansadrS&P 2019 · 1,778 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
Related papers
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
- The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic DatasetsGeorgi Ganev, Emiliano De CristofaroS&P 2025
- PrivSynth: Alternating and Control-Based Optimization for Privacy and Utility in Synthetic DataXinyuan Zhao, Hanlin Gu, Guibao Song, Gongxi Zhu et al.CVPR 2026
- Machine Learning with Privacy for Protected AttributesSaeed Mahloujifar, Chuan Guo, G. Edward Suh, Kamalika ChaudhuriS&P 2025
- On Utility and Privacy in Synthetic Genomic DataBristena Oprisanu, Georgi Ganev, Emiliano De CristofaroNDSS 2022
