Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
Xingli Fang, Jung-Eun Kim
Abstract
Prior approaches for membership privacy preservation usually update or retrain all weights in neural networks, which is costly and can lead to unnecessary utility loss or even more serious misalignment in predictions between training data and non-training data. In this work, we observed three insights: i) privacy vulnerability exists in a very small fraction of weights; ii) however, most of those weights also critically impact utility performance; iii) the importance of weights stems from their locations rather than their values. According to these insights, to preserve privacy, we score critical weights, and instead of discarding those neurons, we rewind only the weights for fine-tuning. We show that, through extensive experiments, this mechanism exhibits outperforming resilience in most cases against Membership Inference Attacks while maintaining utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on38
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
Related papers
- Purifier: Defending Data Inference Attacks via Transforming Confidence ScoresZiqi Yang, Lijin Wang, Da Yang, Jie Wan et al.AAAI 2023 · 20 citations
- Systematic Evaluation of Privacy Risks of Machine Learning ModelsLiwei Song, Prateek MittalUSENIX Security 2021 · 483 citations
- Membership Inference Attacks and Defenses in Neural Network PruningXiaoyong Yuan, Lan ZhangUSENIX Security 2022
- Efficient Privacy Auditing in Federated LearningHongyan Chang, Brandon Edwards, Anindya S. Paul, Reza ShokriUSENIX Security 2024 · 9 citations
- Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?Rui Wen, Michael Backes, Yang ZhangNDSS 2025
