Overconfidence is a Dangerous Thing: Mitigating Membership Inference Attacks by Enforcing Less Confident Prediction
Zitao Chen, Karthik Pattabiraman
摘要
Machine learning (ML) models are vulnerable to membership inference attacks (MIAs), which determine whether a given input is used for training the target model. While there have been many efforts to mitigate MIAs, they often suffer from limited privacy protection, large accuracy drop, and/or requiring additional data that may be difficult to acquire. This work proposes a defense technique, HAMP that can achieve both strong membership privacy and high accuracy, without requiring extra data. To mitigate MIAs in different forms, we observe that they can be unified as they all exploit the ML model's overconfidence in predicting training samples through different proxies. This motivates our design to enforce less confident prediction by the model, hence forcing the model to behave similarly on the training and testing samples. HAMP consists of a novel training framework with high-entropy soft labels and an entropy-based regularizer to constrain the model's prediction while still achieving high accuracy. To further reduce privacy risk, HAMP uniformly modifies all the prediction outputs to become low-confidence outputs while preserving the accuracy, which effectively obscures the differences between the prediction on members and non-members. We conduct extensive evaluation on five benchmark datasets, and show that HAMP provides consistently high accuracy and strong membership privacy. Our comparison with seven state-of-the-art defenses shows that HAMP achieves a superior privacy-utility trade off than those techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 被引用 23 次
- MIST: Defending Against Membership Inference Attacks Through Membership-Invariant Subspace TrainingJiacheng Li, Ninghui Li, Bruno RibeiroUSENIX Security 2024 · 被引用 16 次
- Evaluations of Machine Learning Privacy Defenses are MisleadingMichael Aerni, Jie Zhang, Florian TramèrCCS 2024 · 被引用 12 次
- Mitigating Privacy Risk in Membership Inference by Convex-Concave LossZhenlong Liu, Lei Feng, Huiping Zhuang, Xiaofeng Cao 等ICML 2024 · 被引用 6 次
- Learnability and Privacy Vulnerability are Entangled in a Few Critical WeightsXingli Fang, Jung-Eun KimICLR 2026 · 被引用 1 次
它引用的顶会 Paper26
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
相关 Paper
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 被引用 61 次
- Membership Privacy for Machine Learning Models Through Knowledge TransferVirat Shejwalkar, Amir HoumansadrAAAI 2021 · 被引用 130 次
- Mixup Training for Generative Models to Defend Membership Inference AttacksZhe Ji, Qiansiqi Hu, Liyao Xiang, Chenghu ZhouINFOCOM 2023 · 被引用 3 次
- Mitigating Membership Inference Attacks by Self-Distillation Through a Novel Ensemble ArchitectureXinyu Tang, Saeed Mahloujifar, Liwei Song, Virat Shejwalkar 等USENIX Security 2022
- A Unified Defense Framework Against Membership Inference in Federated Learning via Distillation and Contribution-Aware AggregationLiwei Zhang, Linghui Li, Xiaotian Si, Ziduo Guo 等NDSS 2026 · 被引用 1 次
