Perceptual Adversarial Robustness: Defense Against Unseen Threat Models
Cassidy Laidlaw, Sahil Singla, Soheil Feizi
Abstract
A key challenge in adversarial robustness is the lack of a precise mathematical characterization of human perception, used in the definition of adversarial attacks that are imperceptible to human eyes. Most current attacks and defenses try to avoid this issue by considering restrictive adversarial threat models such as those bounded by L 2 or L ∞ distance, spatial perturbations, etc. However, models that are robust against any of these restrictive threat models are still fragile against other threat models, i.e. they have poor generalization to unforeseen attacks. Moreover, even if a model is robust against the union of several restrictive threat models, it is still susceptible to other imperceptible adversarial examples that are not contained in any of the constituent threat models. To resolve these issues, we propose adversarial training against the set of all imperceptible adversarial examples. Since this set is intractable to compute without a human in the loop, we approximate it using deep neural networks. We call this threat model the neural perceptual threat model (NPTM); it includes adversarial examples with a bounded neural perceptual distance (a neural network-based approximation of the true perceptual distance) to natural images. Through an extensive perceptual study, we show that the neural perceptual distance correlates well with human judgements of perceptibility of adversarial examples, validating our threat model. Under the NPTM, we develop novel perceptual adversarial attacks and defenses. Because the NPTM is very broad, we find that Perceptual Adversarial Training (PAT) against a perceptual attack gives robustness against many other types of adversarial attacks. We test PAT on CIFAR-10 and ImageNet-100 against five diverse adversarial attacks: L 2 , L ∞ , spatial, recoloring, and JPEG. We find that PAT achieves state-of-the-art robustness against the union of these five attacks-more than doubling the accuracy over the next best model-without training against any of them. That is, PAT generalizes well to unforeseen perturbation types. This is vital in sensitive applications where a particular threat model cannot be assumed, and to the best of our knowledge, PAT is the first adversarial training defense with this property.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85c703ae-cd2f-47d8-992d-71fccefc6261Cited by top-tier papers75
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
- Naturalistic Physical Adversarial Patch for Object DetectorsYu-Chih-Tuan Hu, Jun-Cheng Chen, Bo-Han Kung, Kai-Lung Hua et al.ICCV 2021 · 224 citations
- Model-Based Domain GeneralizationAlexander Robey, George J. Pappas, Hamed HassaniNeurIPS 2021 · 167 citations
- Sparse-RS: A Versatile Framework for Query-Efficient Sparse Black-Box Adversarial AttacksFrancesco Croce, Maksym Andriushchenko, Naman D. Singh, Nicolas Flammarion et al.AAAI 2022 · 135 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 171 citations
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 158 citations
Related papers
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
- Sample Efficient Detection and Classification of Adversarial Attacks via Self-Supervised EmbeddingsMazda Moayeri, Soheil FeiziICCV 2021 · 20 citations
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie et al.ICCV 2021 · 35 citations
- Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial AttacksWei-An Lin, Chun Pong Lau, Alexander Levine, Rama Chellappa et al.NeurIPS 2020 · 70 citations
- Provably Adversarially Robust Nearest Prototype ClassifiersVáclav Vorácek, Matthias HeinICML 2022 · 15 citations
