Interpretable Deep Learning under Fire
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, Ting Wang
摘要
Providing explanations for deep neural network (DNN) models is crucial for their use in security-sensitive domains. A plethora of interpretation models have been proposed to help users understand the inner workings of DNNs: how does a DNN arrive at a specific decision for a given input? The improved interpretability is believed to offer a sense of security by involving human in the decision-making process. Yet, due to its data-driven nature, the interpretability itself is potentially susceptible to malicious manipulations, about which little is known thus far. Here we bridge this gap by conducting the first systematic study on the security of interpretable deep learning systems (IDLSes). We show that existing are highly vulnerable to adversarial manipulations. Specifically, we present ADV^2, a new class of attacks that generate adversarial inputs not only misleading target DNNs but also deceiving their coupled interpretation models. Through empirical evaluation against four major types of IDLSes on benchmark datasets and in security-critical applications (e.g., skin cancer diagnosis), we demonstrate that with ADV^2 the adversary is able to arbitrarily designate an input's prediction and interpretation. Further, with both analytical and empirical evidence, we identify the prediction-interpretation gap as one root cause of this vulnerability -- a DNN and its interpretation model are often misaligned, resulting in the possibility of exploiting both models simultaneously. Finally, we explore potential countermeasures against ADV^2, including leveraging its low transferability and incorporating it in an adversarial training framework. Our findings shed light on designing and operating IDLSes in a more secure and informative fashion, leading to several promising research directions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World AttacksYulong Cao, Ningfei Wang, Chaowei Xiao, Dawei Yang 等S&P 2021 · 被引用 309 次
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi 等USENIX Security 2021 · 被引用 241 次
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security ApplicationsDongqi Han, Zhiliang Wang, Wenqi Chen, Ying Zhong 等CCS 2021 · 被引用 108 次
- Interpreting Deep Learning-Based Networking SystemsZili Meng, Minhu Wang, Jiasong Bai, Mingwei Xu 等SIGCOMM 2020 · 被引用 98 次
- Too Good to Be Safe: Tricking Lane Detection in Autonomous Driving with Crafted PerturbationsPengfei Jing, Qiyi Tang, Yuefeng Du, Lei Xue 等USENIX Security 2021 · 被引用 79 次
它引用的顶会 Paper6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov 等S&P 2018 · 被引用 987 次
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su 等CCS 2018 · 被引用 336 次
相关 Paper
- Model-Reuse Attacks on Deep Learning SystemsYujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo 等CCS 2018 · 被引用 197 次
- Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsHengtong Zhang, Jing Gao, Lu SuKDD 2021 · 被引用 21 次
- Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target AttacksBohan Wang, Zewen Liu, Lu Lin, Hui Liu 等ICML 2026 · 被引用 1 次
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao 等KDD 2020 · 被引用 27 次
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary DefenseJingyuan Wang, Yufan Wu, Mingxuan Li, Xin Lin 等KDD 2020 · 被引用 12 次
