Learning to Generate Inversion-Resistant Model Explanations
Hoyong Jeong, Suyoung Lee, Sung Ju Hwang, Sooel Son
Abstract
The wide adoption of deep neural networks (DNNs) in mission-critical applications has spurred the need for interpretable models that provide explanations of the model’s decisions. Unfortunately, previous studies have demonstrated that model explanations facilitate information leakage, rendering DNN models vulnerable to model inversion attacks. These attacks enable the adversary to reconstruct original images based on model explanations, thus leaking privacy-sensitive features. To this end, we present Generative Noise Injector for Model Explanations (GNIME), a novel defense framework that perturbs model explanations to minimize the risk of model inversion attacks while preserving the interpretabilities of the generated explanations. Specifically, we formulate the defense training as a two-player minimax game between the inversion attack network on the one hand, which aims to invert model explanations, and the noise generator network on the other, which aims to inject perturbations to tamper with model inversion attacks. We demonstrate that GNIME significantly decreases the information leakage in model explanations, decreasing transferable classification accuracy in facial recognition models by up to 84.8% while preserving the original functionality of model explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ae732b5-7482-437b-a8d8-58567fd773b1Builds on7
- Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentZiqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai LiangCCS 2019 · 257 citations
- Variational Model Inversion AttacksKuan-Chieh Wang, Yan Fu, Ke Li, Ashish Khisti et al.NeurIPS 2021 · 142 citations
- Knowledge-Enriched Distributional Model Inversion AttacksSi Chen, Mostafa Kahla, Ruoxi Jia, Guo-Jun QiICCV 2021 · 124 citations
- Exploiting Explanations for Model Inversion AttacksXuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. LimICCV 2021 · 113 citations
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
Related papers
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural NetworksYuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang et al.CVPR 2020
- Adversarial Learning of Privacy-Preserving and Task-Oriented RepresentationsTaihong Xiao, Yi-Hsuan Tsai, Kihyuk Sohn, Manmohan Chandraker et al.AAAI 2020 · 87 citations
- InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split InferenceRuijun Deng, Zhihui Lu, Qiang DuanAAAI 2026
- NetGuard: Protecting Commercial Web APIs from Model Inversion Attacks using GAN-generated Fake SamplesXueluan Gong, Ziyao Wang, Yanjiao Chen, Qian Wang et al.WWW 2023 · 8 citations
- Model Inversion Robustness: Can Transfer Learning Help?Sy-Tuyen Ho, Koh Jun Hao, Keshigeyan Chandrasegaran, Ngoc-Bao Nguyen et al.CVPR 2024
