Overconfidence is a Dangerous Thing: Mitigating Membership Inference Attacks by Enforcing Less Confident Prediction
Zitao Chen, Karthik Pattabiraman
Abstract
Machine learning (ML) models are vulnerable to membership inference attacks (MIAs), which determine whether a given input is used for training the target model. While there have been many efforts to mitigate MIAs, they often suffer from limited privacy protection, large accuracy drop, and/or requiring additional data that may be difficult to acquire. This work proposes a defense technique, HAMP that can achieve both strong membership privacy and high accuracy, without requiring extra data. To mitigate MIAs in different forms, we observe that they can be unified as they all exploit the ML model's overconfidence in predicting training samples through different proxies. This motivates our design to enforce less confident prediction by the model, hence forcing the model to behave similarly on the training and testing samples. HAMP consists of a novel training framework with high-entropy soft labels and an entropy-based regularizer to constrain the model's prediction while still achieving high accuracy. To further reduce privacy risk, HAMP uniformly modifies all the prediction outputs to become low-confidence outputs while preserving the accuracy, which effectively obscures the differences between the prediction on members and non-members. We conduct extensive evaluation on five benchmark datasets, and show that HAMP provides consistently high accuracy and strong membership privacy. Our comparison with seven state-of-the-art defenses shows that HAMP achieves a superior privacy-utility trade off than those techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c541f43d-0fb9-487f-ab73-2d6b25b10ea0Cited by top-tier papers15
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 23 citations
- MIST: Defending Against Membership Inference Attacks Through Membership-Invariant Subspace TrainingJiacheng Li, Ninghui Li, Bruno RibeiroUSENIX Security 2024 · 16 citations
- Evaluations of Machine Learning Privacy Defenses are MisleadingMichael Aerni, Jie Zhang, Florian TramèrCCS 2024 · 12 citations
- Mitigating Privacy Risk in Membership Inference by Convex-Concave LossZhenlong Liu, Lei Feng, Huiping Zhuang, Xiaofeng Cao et al.ICML 2024 · 6 citations
- Learnability and Privacy Vulnerability are Entangled in a Few Critical WeightsXingli Fang, Jung-Eun KimICLR 2026 · 1 citation
Builds on26
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
Related papers
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 61 citations
- Membership Privacy for Machine Learning Models Through Knowledge TransferVirat Shejwalkar, Amir HoumansadrAAAI 2021 · 130 citations
- Mixup Training for Generative Models to Defend Membership Inference AttacksZhe Ji, Qiansiqi Hu, Liyao Xiang, Chenghu ZhouINFOCOM 2023 · 3 citations
- Mitigating Membership Inference Attacks by Self-Distillation Through a Novel Ensemble ArchitectureXinyu Tang, Saeed Mahloujifar, Liwei Song, Virat Shejwalkar et al.USENIX Security 2022
- A Unified Defense Framework Against Membership Inference in Federated Learning via Distillation and Contribution-Aware AggregationLiwei Zhang, Linghui Li, Xiaotian Si, Ziduo Guo et al.NDSS 2026 · 1 citation
