SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority Oversampling
Georgi Ganev, MohammadReza Nazari, Rees Davison, Amirhassan Fallah Dizche, XINMIN WU, Ralph Abbey, Jorge G Silva, Emiliano De Cristofaro
Abstract
The Synthetic Minority Over-sampling Technique (SMOTE) is one of the most widely used methods for addressing class imbalance and generating synthetic data. Despite its popularity, little attention has been paid to its privacy implications; yet, it is used in the wild in many privacy-sensitive applications. In this work, we conduct the first systematic study of privacy leakage in SMOTE: We begin by showing that prevailing evaluation practices, i.e., naive distinguishing and distance-to-closest-record metrics, completely fail to detect any leakage and that membership inference attacks (MIAs) can be instantiated with high accuracy. Then, by exploiting SMOTE's geometric properties, we build two novel attacks with very limited assumptions: DistinSMOTE, which perfectly distinguishes real from synthetic records in augmented datasets, and ReconSMOTE, which reconstructs real minority records from synthetic datasets with perfect precision and recall approaching one under realistic imbalance ratios. We also provide theoretical guarantees for both attacks. Experiments on eight standard imbalanced datasets confirm the practicality and effectiveness of these attacks. Overall, our work reveals that SMOTE is inherently non-private and disproportionately exposes minority records, highlighting the need to reconsider its use in privacy-sensitive applications and as a baseline for assessing the privacy of modern generative models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa796f0a-382d-4c6f-893d-889d787fed74Builds on13
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative ModelsDingfan Chen, Ning Yu, Yang Zhang, Mario FritzCCS 2020 · 278 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
Related papers
- The Inadequacy of Similarity-Based Privacy Metrics: Privacy Attacks Against "Truly Anonymous" Synthetic DatasetsGeorgi Ganev, Emiliano De CristofaroS&P 2025
- A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataMeenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc RocherUSENIX Security 2024 · 37 citations
- SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image GenerationYunsung Chung, Yunbei Zhang, Nassir Marrouche, Jihun HammUSENIX Security 2025
- Differential Privacy Under Class Imbalance: Methods and Empirical InsightsLucas Rosenblatt, Yuliia Lut, Ethan Turok, Marco Avella Medina et al.ICML 2025
- On Utility and Privacy in Synthetic Genomic DataBristena Oprisanu, Georgi Ganev, Emiliano De CristofaroNDSS 2022
