Improving Robustness to Model Inversion Attacks via Mutual Information Regularization
Tianhao Wang, Yuheng Zhang, Ruoxi Jia
Abstract
This paper studies defense mechanisms against model inversion (MI) attacks -- a type of privacy attacks aimed at inferring information about the training data distribution given the access to a target machine learning model. Existing defense mechanisms rely on model-specific heuristics or noise injection. While being able to mitigate attacks, existing methods significantly hinder model performance. There remains a question of how to design a defense mechanism that is applicable to a variety of models and achieves better utility-privacy tradeoff.
In this paper, we propose the Mutual Information Regularization based Defense (MID) against MI attacks. The key idea is to limit the information about the model input contained in the prediction, thereby limiting the ability of an adversary to infer the private training attributes from the model prediction. Our defense principle is model-agnostic and we present tractable approximations to the regularizer for linear regression, decision trees, and neural networks, which have been successfully attacked by prior work if not attached with any defenses. We present a formal study of MI attacks by devising a rigorous game-based definition and quantifying the associated information leakage. Our theoretical analysis sheds light on the inefficacy of DP in defending against MI attacks, which has been empirically observed in several prior works. Our experiments demonstrate that MID leads to state-of-the-art performance for a variety of MI attacks, target models and datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d9ba5f9-2ebf-49ac-9970-8dc0abd8579cCited by top-tier papers25
- ResSFL: A Resistance Transfer Framework for Defending Model Inversion Attack in Split Federated LearningJingtao Li, Adnan Siraj Rakin, Xing Chen, Zhezhi He et al.CVPR 2022 · 70 citations
- Label-Only Model Inversion Attacks via Knowledge TransferNgoc-Bao Nguyen, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man CheungNeurIPS 2023 · 44 citations
- Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion AttacksLukas Struppek, Dominik Hintersdorf, Kristian KerstingICLR 2024 · 26 citations
- FedInverse: Evaluating Privacy Leakage in Federated LearningDi Wu, Jun Bai, Yiliao Song, Junjun Chen et al.ICLR 2024 · 22 citations
- Analyzing Privacy Leakage in Machine Learning via Multiple Hypothesis Testing: A Lesson From FanoChuan Guo, Alexandre Sablayrolles, Maziar SanjabiICML 2023 · 20 citations
Builds on3
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Updates-Leak: Data Set Inference and Reconstruction Attacks in Online LearningAhmed Salem, Apratim Bhattacharya, Michael Backes, Mario Fritz et al.USENIX Security 2020
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural NetworksYuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang et al.CVPR 2020
Related papers
- Model Inversion Robustness: Can Transfer Learning Help?Sy-Tuyen Ho, Koh Jun Hao, Keshigeyan Chandrasegaran, Ngoc-Bao Nguyen et al.CVPR 2024
- Trap-MID: Trapdoor-based Defense against Model Inversion AttacksZhenTing Liu, ShangTse ChenNeurIPS 2024 · 12 citations
- Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature FilteringHongyao Yu, Yixiang Qiu, Hao Fang, Tianqu Zhuang et al.KDD 2026 · 2 citations
- Membership Privacy for Machine Learning Models Through Knowledge TransferVirat Shejwalkar, Amir HoumansadrAAAI 2021 · 130 citations
- Bilateral Dependency Optimization: Defending Against Model-inversion AttacksXiong Peng, Feng Liu, Jingfeng Zhang, Long Lan et al.KDD 2022 · 20 citations
