Detecting Adversarial Samples Using Influence Functions and Nearest Neighbors
Gilad Cohen, Guillermo Sapiro, Raja Giryes
摘要
Deep neural networks (DNNs) are notorious for their vulnerability to adversarial attacks, which are small perturbations added to their input images to mislead their prediction. Detection of adversarial examples is, therefore, a fundamental requirement for robust classification frameworks. In this work, we present a method for detecting such adversarial attacks, which is suitable for any pre-trained neural network classifier. We use influence functions to measure the impact of every training sample on the validation set data. From the influence scores, we find the most supportive training samples for any given validation example. A k-nearest neighbor (k-NN) model fitted on the DNN's activation layers is employed to search for the ranking of these supporting training samples. We observe that these samples are highly correlated with the nearest neighbors of the normal inputs, while this correlation is much weaker for adversarial inputs. We train an adversarial detector using the k-NN ranks and distances and show that it successfully distinguishes adversarial examples, getting state-of-the-art results on six attack methods with three datasets. Code is available at https:// github.com/giladcohen/NNIF_adv_defense .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor AttacksJinyuan Jia, Yupei Liu, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2022 · 被引用 90 次
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang 等ICCV 2021 · 被引用 70 次
- Adversarially Robust Conformal PredictionAsaf Gendler, Tsui-Wei Weng, Luca Daniel, Yaniv RomanoICLR 2022 · 被引用 51 次
- Robust Models are less Over-ConfidentJulia Grabinski, Paul Gavrikov, Janis Keuper, Margret KeuperNeurIPS 2022 · 被引用 39 次
- Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context InconsistencyXueping Wang, Shasha Li, Min Liu, Yaonan Wang 等ICCV 2021 · 被引用 34 次
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
相关 Paper
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee 等NDSS 2019 · 被引用 283 次
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang 等AAAI 2020 · 被引用 117 次
- GAT: Generative Adversarial Training for Adversarial Example Detection and Robust ClassificationXuwang Yin, Soheil Kolouri, Gustavo K. RohdeICLR 2020 · 被引用 47 次
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu 等ICML 2024 · 被引用 11 次
- A Study of Defensive Methods to Protect Visual Recommendation Against Adversarial Manipulation of ImagesVito Walter Anelli, Yashar Deldjoo, Tommaso Di Noia, Daniele Malitesta 等SIGIR 2021 · 被引用 30 次
