The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic Datasets
Georgi Ganev, Emiliano De Cristofaro
摘要
Generative models producing synthetic data are meant to provide a privacy-friendly approach to releasing data. However, their privacy guarantees are only considered robust when models satisfy Differential Privacy (DP). Alas, this is not a ubiquitous standard, as many leading companies (and, in fact, research papers) use ad-hoc privacy metrics based on testing the statistical similarity between synthetic and real data. In this paper, we examine the privacy metrics used in real-world synthetic data deployments and demonstrate their unreliability in several ways. First, we provide counter-examples where severe privacy violations occur even if the privacy tests pass and instantiate accurate membership and attribute inference attacks with minimal cost. We then introduce Recon-Syn, a reconstruction attack that generates multiple synthetic datasets that are considered private by the metrics but actually leak information unique to individual records. We show that ReconSyn recovers 78-100% of the outliers in the train data with only black-box access to a single fitted generative model and the privacy metrics. In the process, we show that applying DP only to the model does not mitigate this attack, as using privacy metrics breaks the end-to-end DP pipeline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority OversamplingGeorgi Ganev, MohammadReza Nazari, Rees Davison, Amirhassan Fallah Dizche 等ICLR 2026 · 被引用 6 次
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 被引用 2 次
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang 等ICML 2026
它引用的顶会 Paper29
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
相关 Paper
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 被引用 37 次
- Bounding Training Data Reconstruction in Private (Deep) LearningChuan Guo, Brian Karrer, Kamalika Chaudhuri, Laurens van der MaatenICML 2022 · 被引用 66 次
- "What do you want from theory alone?" Experimenting with Tight Auditing of Differentially Private Synthetic Data GenerationMeenatchi Sundaram Muthu Selva Annamalai, Georgi Ganev, Emiliano De CristofaroUSENIX Security 2024 · 被引用 24 次
- Synthesizing Plausible Privacy-Preserving Location TracesVincent Bindschaedler, Reza ShokriS&P 2016 · 被引用 193 次
- Understanding Disclosure Risk in Differential Privacy with Applications to Noise Calibration and AuditingPatricia Guerra-Balboa, Annika Sauer, Héber Hwang Arcolezi, Thorsten StrufeVLDB 2026
