Exploiting Explanations for Model Inversion Attacks
Xuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. Lim
摘要
The successful deployment of artificial intelligence (AI) in many domains from healthcare to hiring requires their responsible use, particularly in model explanations and privacy. Explainable artificial intelligence (XAI) provides more information to help users to understand model decisions, yet this additional knowledge exposes additional risks for privacy attacks. Hence, providing explanation harms privacy. We study this risk for image-based model inversion attacks and identified several attack architectures with increasing performance to reconstruct private image data from model explanations. We have developed several multi-modal transposed CNN architectures that achieve significantly higher inversion performance than using the target model prediction only. These XAI-aware inversion models were designed to exploit the spatial knowledge in image explanations. To understand which explanations have higher privacy risk, we analyzed how various explanation types and factors influence inversion performance. In spite of some models not providing explanations, we further demonstrate increased inversion performance even for non-explainable target models by exploiting explanations of surrogate models through attention transfer. This method first inverts an explanation from the target prediction, then reconstructs the target image. These threats highlight the urgent and significant privacy risks of explanations and calls attention for new privacy preservation techniques that balance the dual-requirement for AI explainability and privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Plug & Play Attacks: Towards Robust and Flexible Model Inversion AttacksLukas Struppek, Dominik Hintersdorf, Antonio De Almeida Correia, Antonia Adler 等ICML 2022 · 被引用 88 次
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 被引用 76 次
- On Strengthening and Defending Graph Reconstruction Attack with Markov Chain ApproximationZhanke Zhou, Chenyu Zhou, Xuan Li, Jiangchao Yao 等ICML 2023 · 被引用 25 次
- XRand: Differentially Private Defense against Explanation-Guided AttacksTruc D. T. Nguyen, Phung Lai, Hai Phan, My T. ThaiAAAI 2023 · 被引用 22 次
- Feature Inference Attack on Shapley ValuesXinjian Luo, Yangfan Jiang, Xiaokui XiaoCCS 2022 · 被引用 20 次
它引用的顶会 Paper8
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- MemGuard: Defending against Black-Box Membership Inference Attacks via Adversarial ExamplesJinyuan Jia, Ahmed Salem, Michael Backes, Yang Zhang 等CCS 2019 · 被引用 464 次
- Neural Network Inversion in Adversarial Setting via Background Knowledge AlignmentZiqi Yang, Jiyi Zhang, Ee-Chien Chang, Zhenkai LiangCCS 2019 · 被引用 257 次
相关 Paper
- Please Tell Me More: Privacy Impact of Explainability through the Lens of Membership Inference AttackHan Liu, Yuhao Wu, Zhiyuan Yu, Ning ZhangS&P 2024 · 被引用 52 次
- Learning to Generate Inversion-Resistant Model ExplanationsHoyong Jeong, Suyoung Lee, Sung Ju Hwang, Sooel SonNeurIPS 2022 · 被引用 4 次
- Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model QueriesYu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin 等ACL 2024 · 被引用 5 次
- Adversarial Learning of Privacy-Preserving and Task-Oriented RepresentationsTaihong Xiao, Yi-Hsuan Tsai, Kihyuk Sohn, Manmohan Chandraker 等AAAI 2020 · 被引用 87 次
- Variational Model Inversion AttacksKuan-Chieh Wang, Yan Fu, Ke Li, Ashish Khisti 等NeurIPS 2021 · 被引用 142 次
