When Does Data Augmentation Help With Membership Inference Attacks?
Yigitcan Kaya, Tudor Dumitras
Abstract
Deep learning models often raise privacy concerns as they leak information about their training data. This leakage enables membership inference attacks (MIA) that can identify whether a data point was in a model's training set. Research shows that some data augmentation mechanisms may reduce the risk by combatting a key factor increasing the leakage, overfitting. While many mechanisms exist, their effectiveness against MIAs and privacy properties have not been studied systematically. Employing two recent MIAs, we explore the lower bound on the risk in the absence of formal upper bounds. First, we evaluate 7 mechanisms and differential privacy, on three image classification tasks. We find that applying augmentation to increase the model's utility does not mitigate the risk and protection comes with a utility penalty. Further, we also investigate why popular label smoothing mechanism consistently amplifies the risk. Finally, we propose loss-rank-correlation (LRC) metric to assess how similar the effects of different mechanisms are. This, for example, reveals the similarity of applying high-intensity augmentation against MIAs to simply reducing the training time. Our findings emphasize the utilityprivacy trade-off and provide practical guidelines on using augmentation to manage the trade-off.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6f0786c-5d9d-4162-9ac6-d239b0a086cdCited by top-tier papers17
- Bag of Tricks for Training Data Extraction from Language ModelsWeichen Yu, Tianyu Pang, Qian Liu, Chao Du et al.ICML 2023 · 87 citations
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 61 citations
- Practical Membership Inference Attacks Against Large-Scale Multi-Modal Models: A Pilot StudyMyeongseob Ko, Ming Jin, Chenguang Wang, Ruoxi JiaICCV 2023 · 51 citations
- Parameters or Privacy: A Provable Tradeoff Between Overparameterization and Membership InferenceJasper Tan, Blake Mason, Hamid Javadi, Richard G. BaraniukNeurIPS 2022 · 22 citations
- On the Privacy Risks of Cell-Based NAS ArchitecturesHai Huang, Zhikun Zhang, Yun Shen, Michael Backes et al.CCS 2022 · 6 citations
Builds on6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 586 citations
- How Does Data Augmentation Affect Privacy in Machine Learning?Da Yu, Huishuai Zhang, Wei Chen, Jian Yin et al.AAAI 2021 · 67 citations
Related papers
- Mitigating Privacy Risk in Membership Inference by Convex-Concave LossZhenlong Liu, Lei Feng, Huiping Zhuang, Xiaofeng Cao et al.ICML 2024 · 6 citations
- Mixup Training for Generative Models to Defend Membership Inference AttacksZhe Ji, Qiansiqi Hu, Liyao Xiang, Chenghu ZhouINFOCOM 2023 · 3 citations
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 628 citations
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 543 citations
- Free Record-Level Privacy Risk Evaluation Through Artifact-Based MethodsJoseph Pollock, Igor Shilov, Euodia Dodd, Yves-Alexandre de MontjoyeUSENIX Security 2025
