USENIX Security2020Top-tier venue
Interpretable Deep Learning under Fire
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, Ting Wang
Abstract
Providing explanations for deep neural network (DNN) models is crucial for their use in security-sensitive domains. A plethora of interpretation models have been proposed to help users understand the inner workings of DNNs: how does a DNN arrive at a specific decision for a given input? The improved interpretability is believed to offer a sense of security by involving human in the decision-making process. Yet, due to its data-driven nature, the interpretability itself is potentially susceptible to malicious manipulations, about which little is known thus far. Here we bridge this gap by conducting the first systematic study on the security of interpretable deep learning systems (IDLSes). We show that existing are highly vulnerable to adversarial manipulations. Specifically, we present ADV^2, a new class of attacks that generate adversarial inputs not only misleading target DNNs but also deceiving their coupled interpretation models. Through empirical evaluation against four major types of IDLSes on benchmark datasets and in security-critical applications (e.g., skin cancer diagnosis), we demonstrate that with ADV^2 the adversary is able to arbitrarily designate an input's prediction and interpretation. Further, with both analytical and empirical evidence, we identify the prediction-interpretation gap as one root cause of this vulnerability -- a DNN and its interpretation model are often misaligned, resulting in the possibility of exploiting both models simultaneously. Finally, we explore potential countermeasures against ADV^2, including leveraging its low transferability and incorporating it in an adversarial training framework. Our findings shed light on designing and operating IDLSes in a more secure and informative fashion, leading to several promising research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc312669-2c40-4216-ae4c-b180f8121a3dCited by top-tier papers47
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World AttacksYulong Cao, Ningfei Wang, Chaowei Xiao, Dawei Yang et al.S&P 2021 · 309 citations
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security ApplicationsDongqi Han, Zhiliang Wang, Wenqi Chen, Ying Zhong et al.CCS 2021 · 108 citations
- Interpreting Deep Learning-Based Networking SystemsZili Meng, Minhu Wang, Jiasong Bai, Mingwei Xu et al.SIGCOMM 2020 · 98 citations
- Too Good to Be Safe: Tricking Lane Detection in Autonomous Driving with Crafted PerturbationsPengfei Jing, Qiyi Tang, Yuefeng Du, Lei Xue et al.USENIX Security 2021 · 79 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
Related papers
- Model-Reuse Attacks on Deep Learning SystemsYujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo et al.CCS 2018 · 197 citations
- Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsHengtong Zhang, Jing Gao, Lu SuKDD 2021 · 21 citations
- Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target AttacksBohan Wang, Zewen Liu, Lu Lin, Hui Liu et al.ICML 2026 · 1 citation
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao et al.KDD 2020 · 27 citations
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary DefenseJingyuan Wang, Yufan Wu, Mingxuan Li, Xin Lin et al.KDD 2020 · 12 citations
