Students Parrot Their Teachers: Membership Inference on Model Distillation
Matthew Jagielski, Milad Nasr, Katherine Lee, Christopher A. Choquette-Choo, Nicholas Carlini, Florian Tramèr
Abstract
Model distillation is frequently proposed as a technique to reduce the privacy leakage of machine learning. These empirical privacy defenses rely on the intuition that distilled student'' models protect the privacy of training data, as they only interact with this data indirectly through a teacher'' model. In this work, we design membership inference attacks to systematically study the privacy provided by knowledge distillation to both the teacher and student training sets. Our new attacks show that distillation alone provides only limited privacy across a number of domains. We explain the success of our attacks on distillation by showing that membership inference attacks on a private dataset can succeed even if the target model is never queried on any actual training points, but only on inputs whose predictions are highly influenced by training data. Finally, we show that our attacks are strongest when student and teacher sets are similar, or when the attacker can poison the teacher set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77e54bcf-4f19-4b6b-8f33-e21f71ab187cCited by top-tier papers13
- Teach LLMs to Phish: Stealing Private Information from Language ModelsAshwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang et al.ICLR 2024 · 41 citations
- Auditing Private PredictionKaran Chadha, Matthew Jagielski, Nicolas Papernot, Christopher A. Choquette-Choo et al.ICML 2024 · 10 citations
- Cascading and Proxy Membership Inference AttacksYuntao Du, Jiacheng Li, Yuetian Chen, Kaiyuan Zhang et al.NDSS 2026 · 8 citations
- Generalizing Trust: Weak-to-Strong Trustworthiness in Language ModelsLillian Sun, Martin Pawelczyk, Zhenting Qi, Aounon Kumar et al.ACL 2026 · 7 citations
- S-RAG: A Novel Audit Framework for Detecting Unauthorized Use of Personal Data in RAG SystemsZhirui Zeng, Jiamou Liu, Meng-Fen Chiang, Jialing He et al.ACL 2025 · 4 citations
Builds on19
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 628 citations
Related papers
- Membership Privacy for Machine Learning Models Through Knowledge TransferVirat Shejwalkar, Amir HoumansadrAAAI 2021 · 130 citations
- Mitigating Membership Inference Attacks by Self-Distillation Through a Novel Ensemble ArchitectureXinyu Tang, Saeed Mahloujifar, Liwei Song, Virat Shejwalkar et al.USENIX Security 2022
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 543 citations
- Membership Inference Attacks by Exploiting Loss TrajectoryYiyong Liu, Zhengyu Zhao, Michael Backes, Yang ZhangCCS 2022 · 79 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
