Detecting Adversarial Samples Using Influence Functions and Nearest Neighbors
Gilad Cohen, Guillermo Sapiro, Raja Giryes
Abstract
Deep neural networks (DNNs) are notorious for their vulnerability to adversarial attacks, which are small perturbations added to their input images to mislead their prediction. Detection of adversarial examples is, therefore, a fundamental requirement for robust classification frameworks. In this work, we present a method for detecting such adversarial attacks, which is suitable for any pre-trained neural network classifier. We use influence functions to measure the impact of every training sample on the validation set data. From the influence scores, we find the most supportive training samples for any given validation example. A k-nearest neighbor (k-NN) model fitted on the DNN's activation layers is employed to search for the ranking of these supporting training samples. We observe that these samples are highly correlated with the nearest neighbors of the normal inputs, while this correlation is much weaker for adversarial inputs. We train an adversarial detector using the k-NN ranks and distances and show that it successfully distinguishes adversarial examples, getting state-of-the-art results on six attack methods with three datasets. Code is available at https:// github.com/giladcohen/NNIF_adv_defense .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef1971b8-cce8-4261-89be-95eec4567426Cited by top-tier papers17
- Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor AttacksJinyuan Jia, Yupei Liu, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2022 · 90 citations
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang et al.ICCV 2021 · 70 citations
- Adversarially Robust Conformal PredictionAsaf Gendler, Tsui-Wei Weng, Luca Daniel, Yaniv RomanoICLR 2022 · 51 citations
- Robust Models are less Over-ConfidentJulia Grabinski, Paul Gavrikov, Janis Keuper, Margret KeuperNeurIPS 2022 · 39 citations
- Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context InconsistencyXueping Wang, Shasha Li, Min Liu, Yaonan Wang et al.ICCV 2021 · 34 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
Related papers
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee et al.NDSS 2019 · 283 citations
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang et al.AAAI 2020 · 117 citations
- GAT: Generative Adversarial Training for Adversarial Example Detection and Robust ClassificationXuwang Yin, Soheil Kolouri, Gustavo K. RohdeICLR 2020 · 47 citations
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu et al.ICML 2024 · 11 citations
- A Study of Defensive Methods to Protect Visual Recommendation Against Adversarial Manipulation of ImagesVito Walter Anelli, Yashar Deldjoo, Tommaso Di Noia, Daniele Malitesta et al.SIGIR 2021 · 30 citations
