NIC: Detecting Adversarial Samples with Neural Network Invariant Checking
Shiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee, Xiangyu Zhang
Abstract
Deep Neural Networks (DNN) are vulnerable to adversarial samples that are generated by perturbing correctly classified inputs to cause DNN models to misbehave (e.g., misclassification). This can potentially lead to disastrous consequences especially in security-sensitive applications. Existing defense and detection techniques work well for specific attacks under various assumptions (e.g., the set of possible attacks are known beforehand). However, they are not sufficiently general to protect against a broader range of attacks. In this paper, we analyze the internals of DNN models under various attacks and identify two common exploitation channels: the provenance channel and the activation value distribution channel. We then propose a novel technique to extract DNN invariants and use them to perform runtime adversarial sample detection. Our experimental results of 11 different kinds of attacks on popular datasets including ImageNet and 13 models show that our technique can effectively detect all these attacks (over 90% accuracy) with limited false positives. We also compare it with three state-of-theart techniques including the Local Intrinsic Dimensionality (LID) based method, denoiser based methods (i.e., MagNet and HGD), and the prediction inconsistency based approach (i.e., feature squeezing). Our experiments show promising results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers53
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma et al.CCS 2019 · 531 citations
- Adversarial Neuron Pruning Purifies Backdoored Deep ModelsDongxian Wu, Yisen WangNeurIPS 2021 · 441 citations
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov et al.S&P 2021 · 381 citations
- Rethinking the Backdoor Attacks' Triggers: A Frequency PerspectiveYi Zeng, Won Park, Z. Morley Mao, Ruoxi JiaICCV 2021 · 274 citations
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu et al.FSE 2020 · 165 citations
Builds on16
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 1,581 citations
Related papers
- Detecting Adversarial Samples Using Influence Functions and Nearest NeighborsGilad Cohen, Guillermo Sapiro, Raja GiryesCVPR 2020
- Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform DomainJinyu Tian, Jiantao Zhou, Yuanman Li, Jia DuanAAAI 2021 · 72 citations
- DNNGuard: An Elastic Heterogeneous DNN Accelerator Architecture against Adversarial AttacksXingbin Wang, Rui Hou, Boyan Zhao, Fengkai Yuan et al.ASPLOS 2020 · 33 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
