Machine Learning Models that Remember Too Much
Congzheng Song, Thomas Ristenpart, Vitaly Shmatikov
Abstract
Machine learning (ML) is becoming a commodity. Numerous ML frameworks and services are available to data holders who are not ML experts but want to train predictive models on their data. It is important that ML models trained on sensitive inputs (e.g., personal images or documents) not leak too much information about the training data. We consider a malicious ML provider who supplies model-training code to the data holder, does not observe the training, but then obtains white- or black-box access to the resulting model. In this setting, we design and implement practical algorithms, some of them very similar to standard ML techniques such as regularization and data augmentation, that "memorize" information about the training dataset in the modelyet the model is as accurate and predictive as a conventionally trained model. We then explain how the adversary can extract memorized information from the model. We evaluate our techniques on standard ML tasks for image classification (CIFAR10), face recognition (LFW and FaceScrub), and text analysis (20 Newsgroups and IMDB). In all cases, we show how our algorithms create models that have high predictive power yet allow accurate extraction of subsets of their training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc0f0f8b-5739-4516-a12e-0afe60139270Cited by top-tier papers80
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Property Inference Attacks on Fully Connected Neural Networks using Permutation Invariant RepresentationsKaran Ganju, Qi Wang, Wei Yang, Carl A. Gunter et al.CCS 2018 · 574 citations
- Stealing Hyperparameters in Machine LearningBinghui Wang, Neil Zhenqiang GongS&P 2018 · 504 citations
Builds on5
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Oblivious Multi-Party Machine Learning on Trusted ProcessorsOlga Ohrimenko, Felix Schuster, Cédric Fournet, Aastha Mehta et al.USENIX Security 2016 · 594 citations
- Membership Privacy in MicroRNA-based StudiesMichael Backes, Pascal Berrang, Mathias Humbert, Praveen ManoharanCCS 2016 · 154 citations
- On Omitting Commits and Committing Omissions: Preventing Git Metadata Tampering That (Re)introduces Software VulnerabilitiesSantiago Torres-Arias, Anil Kumar Ammula, Reza Curtmola, Justin CapposUSENIX Security 2016 · 33 citations
Related papers
- A Method to Facilitate Membership Inference Attacks in Deep Learning ModelsZitao Chen, Karthik PattabiramanNDSS 2025
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Stealing Your Data from Compressed Machine Learning ModelsNuo Xu, Qi Liu, Tao Liu, Zihao Liu et al.DAC 2020 · 3 citations
- Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentZiqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai LiangCCS 2019 · 257 citations
- Reconstructing Training Data with Informed AdversariesBorja Balle, Giovanni Cherubin, Jamie HayesS&P 2022 · 214 citations
