Membership Inference Attacks From First Principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, Florian Tramèr
Abstract
A membership inference attack allows an adversary to query a trained machine learning model to predict whether or not a particular example was contained in the model’s training dataset. These attacks are currently evaluated using average-case “accuracy” metrics that fail to characterize whether the attack can confidently identify any members of the training set. We argue that attacks should instead be evaluated by computing their true-positive rate at low (e.g., ≤ 0.1%) false-positive rates, and find most prior attacks perform poorly when evaluated in this way. To address this we develop a Likelihood Ratio Attack (LiRA) that carefully combines multiple ideas from the literature. Our attack is more powerful at low false-positive rates, and also strictly dominates prior attacks on existing metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ecf16a9-3510-4d02-ae87-4c823f6d1ba1Cited by top-tier papers353
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 363 citations
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong et al.ICLR 2024 · 351 citations
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 217 citations
Builds on25
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated LearningMilad Nasr, Reza Shokri, Amir HoumansadrS&P 2019 · 1,778 citations
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
Related papers
- Low-Cost High-Power Membership Inference AttacksSajjad Zarifzadeh, Philippe Liu, Reza ShokriICML 2024 · 92 citations
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2025 · 26 citations
- Scalable Membership Inference Attacks via Quantile RegressionMartin Bertran Lopez, Shuai Tang, Aaron Roth, Michael Kearns et al.NeurIPS 2023 · 96 citations
- Is Difficulty Calibration All We Need? Towards More Practical Membership Inference AttacksYu He, Boheng Li, Yao Wang, Mengda Yang et al.CCS 2024 · 4 citations
- Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference AttackJing Xue, Zhishen Sun, Haishan Ye, Luo Luo et al.AAAI 2026
