Beating Attackers At Their Own Games: Adversarial Example Detection Using Adversarial Gradient Directions
Yuhang Wu, Sunpreet S. Arora, Yanhong Wu, Hao Yang
摘要
Adversarial examples are input examples that are specifically crafted to deceive machine learning classifiers. State-of-the-art adversarial example detection methods characterize an input example as adversarial either by quantifying the magnitude of feature variations under multiple perturbations or by measuring its distance from estimated benign example distribution. Instead of using such metrics, the proposed method is based on the observation that the directions of adversarial gradients when crafting (new) adversarial examples play a key role in characterizing the adversarial space. Compared to detection methods that use multiple perturbations, the proposed method is efficient as it only applies a single random perturbation on the input example. Experiments conducted on two different databases, CIFAR-10 and ImageNet, show that the proposed detection method achieves, respectively, 97.9% and 98.6% AUC-ROC (on average) on five different adversarial attacks, and outperforms multiple state-of-the-art detection methods. Results demonstrate the effectiveness of using adversarial gradient directions for adversarial example detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- ML-LOO: Detecting Adversarial Examples with Feature AttributionPuyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang 等AAAI 2020 · 被引用 117 次
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang 等ICCV 2021 · 被引用 70 次
- Improving Adversarial Transferability via Intermediate-level Perturbation DecayQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2023 · 被引用 43 次
- ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty EstimationFan Yin, Yao Li, Cho-Jui Hsieh, Kai-Wei ChangEMNLP 2022 · 被引用 6 次
- Adversarial Training for Raw-Binary Malware ClassifiersKeane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer 等USENIX Security 2023
