When Does Data Augmentation Help With Membership Inference Attacks?
Yigitcan Kaya, Tudor Dumitras
摘要
Deep learning models often raise privacy concerns as they leak information about their training data. This leakage enables membership inference attacks (MIA) that can identify whether a data point was in a model's training set. Research shows that some data augmentation mechanisms may reduce the risk by combatting a key factor increasing the leakage, overfitting. While many mechanisms exist, their effectiveness against MIAs and privacy properties have not been studied systematically. Employing two recent MIAs, we explore the lower bound on the risk in the absence of formal upper bounds. First, we evaluate 7 mechanisms and differential privacy, on three image classification tasks. We find that applying augmentation to increase the model's utility does not mitigate the risk and protection comes with a utility penalty. Further, we also investigate why popular label smoothing mechanism consistently amplifies the risk. Finally, we propose loss-rank-correlation (LRC) metric to assess how similar the effects of different mechanisms are. This, for example, reveals the similarity of applying high-intensity augmentation against MIAs to simply reducing the training time. Our findings emphasize the utilityprivacy trade-off and provide practical guidelines on using augmentation to manage the trade-off.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Bag of Tricks for Training Data Extraction from Language ModelsWeichen Yu, Tianyu Pang, Qian Liu, Chao Du 等ICML 2023 · 被引用 87 次
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 被引用 61 次
- Practical Membership Inference Attacks Against Large-Scale Multi-Modal Models: A Pilot StudyMyeongseob Ko, Ming Jin, Chenguang Wang, Ruoxi JiaICCV 2023 · 被引用 51 次
- Parameters or Privacy: A Provable Tradeoff Between Overparameterization and Membership InferenceJasper Tan, Blake Mason, Hamid Javadi, Richard G. BaraniukNeurIPS 2022 · 被引用 22 次
- On the Privacy Risks of Cell-Based NAS ArchitecturesHai Huang, Zhikun Zhang, Yun Shen, Michael Backes 等CCS 2022 · 被引用 6 次
它引用的顶会 Paper6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 被引用 586 次
- How Does Data Augmentation Affect Privacy in Machine Learning?Da Yu, Huishuai Zhang, Wei Chen, Jian Yin 等AAAI 2021 · 被引用 67 次
相关 Paper
- Mitigating Privacy Risk in Membership Inference by Convex-Concave LossZhenlong Liu, Lei Feng, Huiping Zhuang, Xiaofeng Cao 等ICML 2024 · 被引用 6 次
- Mixup Training for Generative Models to Defend Membership Inference AttacksZhe Ji, Qiansiqi Hu, Liyao Xiang, Chenghu ZhouINFOCOM 2023 · 被引用 3 次
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 被引用 628 次
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 被引用 543 次
- Free Record-Level Privacy Risk Evaluation Through Artifact-Based MethodsJoseph Pollock, Igor Shilov, Euodia Dodd, Yves-Alexandre de MontjoyeUSENIX Security 2025
