Beating Attackers At Their Own Games: Adversarial Example Detection Using Adversarial Gradient Directions
Yuhang Wu, Sunpreet S. Arora, Yanhong Wu, Hao Yang
Abstract
Adversarial examples are input examples that are specifically crafted to deceive machine learning classifiers. State-of-the-art adversarial example detection methods characterize an input example as adversarial either by quantifying the magnitude of feature variations under multiple perturbations or by measuring its distance from estimated benign example distribution. Instead of using such metrics, the proposed method is based on the observation that the directions of adversarial gradients when crafting (new) adversarial examples play a key role in characterizing the adversarial space. Compared to detection methods that use multiple perturbations, the proposed method is efficient as it only applies a single random perturbation on the input example. Experiments conducted on two different databases, CIFAR-10 and ImageNet, show that the proposed detection method achieves, respectively, 97.9% and 98.6% AUC-ROC (on average) on five different adversarial attacks, and outperforms multiple state-of-the-art detection methods. Results demonstrate the effectiveness of using adversarial gradient directions for adversarial example detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2db0bba-118a-4884-9054-df78ea2a2e33Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang et al.AAAI 2020 · 117 citations
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang et al.ICCV 2021 · 70 citations
- Improving Adversarial Transferability via Intermediate-level Perturbation DecayQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2023 · 43 citations
- ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty EstimationFan Yin, Yao Li, Cho-Jui Hsieh, Kai-Wei ChangEMNLP 2022 · 6 citations
- Adversarial Training for Raw-Binary Malware ClassifiersKeane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer et al.USENIX Security 2023
