ML-LOO: Detecting Adversarial Examples with Feature Attribution
Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang, Michael I. Jordan
Abstract
Deep neural networks obtain state-of-the-art performance on a series of tasks. However, they are easily fooled by adding a small adversarial perturbation to the input. The perturbation is often imperceptible to humans on image data. We observe a significant difference in feature attributions between adversarially crafted examples and original examples. Based on this observation, we introduce a new framework to detect adversarial examples through thresholding a scale estimate of feature attribution scores. Furthermore, we extend our method to include multi-layer feature attributions in order to tackle attacks that have mixed confidence levels. As demonstrated in extensive experiments, our method achieves superior performances in distinguishing adversarial examples from popular attack methods on a variety of real data sets compared to state-of-the-art detection methods. In particular, our method is able to detect adversarial examples of mixed confidence levels, and transfer between different attacking methods. We also show that our method achieves competitive performance even when the attacker has complete access to the detector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Towards a Unified Game-Theoretic View of Adversarial Perturbations and RobustnessJie Ren, Die Zhang, Yisen Wang, Lu Chen et al.NeurIPS 2021 · 27 citations
- Reverse Engineering of Imperceptible Adversarial Image PerturbationsYifan Gong, Yuguang Yao, Yize Li, Yimeng Zhang et al.ICLR 2022 · 25 citations
- Increasing Confidence in Adversarial Robustness EvaluationsRoland S. Zimmermann, Wieland Brendel, Florian Tramèr, Nicholas CarliniNeurIPS 2022 · 25 citations
- SPADE: A Spectral Method for Black-Box Adversarial Robustness EvaluationWuxinlin Cheng, Chenhui Deng, Zhiqiang Zhao, Yaohui Cai et al.ICML 2021 · 24 citations
- Synergy-of-Experts: Collaborate to Improve Adversarial RobustnessSen Cui, Jingfeng Zhang, Jian Liang, Bo Han et al.NeurIPS 2022 · 12 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
Related papers
- Detecting Adversarial Samples Using Influence Functions and Nearest NeighborsGilad Cohen, Guillermo Sapiro, Raja GiryesCVPR 2020
- Beating Attackers At Their Own Games: Adversarial Example Detection Using Adversarial Gradient DirectionsYuhang Wu, Sunpreet S. Arora, Yanhong Wu, Hao YangAAAI 2021 · 12 citations
- Introducing Competition to Boost the Transferability of Targeted Adversarial Examples Through Clean Feature MixupJunyoung Byun, Myung-Joon Kwon, Seungju Cho, Yoonji Kim et al.CVPR 2023
- Adversarial Camouflage: Hiding Physical-World Attacks With Natural StylesRanjie Duan, Xingjun Ma, Yisen Wang, James Bailey et al.CVPR 2020
- Attack as defense: characterizing adversarial examples using robustnessZhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang et al.ISSTA 2021 · 34 citations
