On the Importance of Difficulty Calibration in Membership Inference Attacks
Lauren Watson, Chuan Guo, Graham Cormode, Alexandre Sablayrolles
Abstract
The vulnerability of machine learning models to membership inference attacks has received much attention in recent years. However, existing attacks mostly remain impractical due to having high false positive rates, where non-member samples are often erroneously predicted as members. This type of error makes the predicted membership signal unreliable, especially since most samples are non-members in real world applications. In this work, we argue that membership inference attacks can benefit drastically from difficulty calibration, where an attack's predicted membership score is adjusted to the difficulty of correctly classifying the target sample. We show that difficulty calibration can significantly reduce the false positive rate of a variety of existing attacks without a loss in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc3ba3de-3de8-465c-ae4e-40d706d52442Cited by top-tier papers71
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Enhanced Membership Inference Attacks against Machine Learning ModelsJiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler et al.CCS 2022 · 150 citations
- Scalable Membership Inference Attacks via Quantile RegressionMartin Bertran Lopez, Shuai Tang, Aaron Roth, Michael Kearns et al.NeurIPS 2023 · 96 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated LearningMilad Nasr, Reza Shokri, Amir HoumansadrS&P 2019 · 1,778 citations
Related papers
- Is Difficulty Calibration All We Need? Towards More Practical Membership Inference AttacksYu He, Boheng Li, Yao Wang, Mengda Yang et al.CCS 2024 · 4 citations
- Chameleon: Increasing Label-Only Membership Leakage with Adaptive PoisoningHarsh Chaudhari, Giorgio Severi, Alina Oprea, Jonathan R. UllmanICLR 2024 · 8 citations
- On the Difficulty of Membership Inference AttacksShahbaz Rezaei, Xin LiuCVPR 2021
- MemGuard: Defending against Black-Box Membership Inference Attacks via Adversarial ExamplesJinyuan Jia, Ahmed Salem, Michael Backes, Yang Zhang et al.CCS 2019 · 464 citations
- You Only Query Once: An Efficient Label-Only Membership Inference AttackYutong Wu, Han Qiu, Shangwei Guo, Jiwei Li et al.ICLR 2024 · 20 citations
