Learning to Generate Inversion-Resistant Model Explanations
Hoyong Jeong, Suyoung Lee, Sung Ju Hwang, Sooel Son
摘要
The wide adoption of deep neural networks (DNNs) in mission-critical applications has spurred the need for interpretable models that provide explanations of the model’s decisions. Unfortunately, previous studies have demonstrated that model explanations facilitate information leakage, rendering DNN models vulnerable to model inversion attacks. These attacks enable the adversary to reconstruct original images based on model explanations, thus leaking privacy-sensitive features. To this end, we present Generative Noise Injector for Model Explanations (GNIME), a novel defense framework that perturbs model explanations to minimize the risk of model inversion attacks while preserving the interpretabilities of the generated explanations. Specifically, we formulate the defense training as a two-player minimax game between the inversion attack network on the one hand, which aims to invert model explanations, and the noise generator network on the other, which aims to inject perturbations to tamper with model inversion attacks. We demonstrate that GNIME significantly decreases the information leakage in model explanations, decreasing transferable classification accuracy in facial recognition models by up to 84.8% while preserving the original functionality of model explanations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentZiqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai LiangCCS 2019 · 被引用 257 次
- Variational Model Inversion AttacksKuan-Chieh Wang, Yan Fu, Ke Li, Ashish Khisti 等NeurIPS 2021 · 被引用 142 次
- Knowledge-Enriched Distributional Model Inversion AttacksSi Chen, Mostafa Kahla, Ruoxi Jia, Guo-Jun QiICCV 2021 · 被引用 124 次
- Exploiting Explanations for Model Inversion AttacksXuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. LimICCV 2021 · 被引用 113 次
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 被引用 22 次
相关 Paper
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural NetworksYuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang 等CVPR 2020
- Adversarial Learning of Privacy-Preserving and Task-Oriented RepresentationsTaihong Xiao, Yi-Hsuan Tsai, Kihyuk Sohn, Manmohan Chandraker 等AAAI 2020 · 被引用 87 次
- InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split InferenceRuijun Deng, Zhihui Lu, Qiang DuanAAAI 2026
- NetGuard: Protecting Commercial Web APIs from Model Inversion Attacks using GAN-generated Fake SamplesXueluan Gong, Ziyao Wang, Yanjiao Chen, Qian Wang 等WWW 2023 · 被引用 8 次
- Model Inversion Robustness: Can Transfer Learning Help?Sy-Tuyen Ho, Koh Jun Hao, Keshigeyan Chandrasegaran, Ngoc-Bao Nguyen 等CVPR 2024
