Doppelgangers and Adversarial Vulnerability
George I. Kamberov
Abstract
Machine learning (ML) classifiers can make mistakes that are perceptually and cognitively disturbing to humans. The most notorious examples of such errors are adversarial visual metamers. This paper investigates the phenomenon of adversarial Doppelgängers (AD), which encompasses adversarial visual metamers, and compares the performance and robustness of ML classifiers to human performance. We find that ADs are inputs that are close to each other with respect to a perceptual metric defined in this paper, and show that ADs are qualitatively different from the usual adversarial examples. The vast majority of classifiers are vulnerable to ADs and robustness-accuracy trade-offs may not improve them. Some classification problems do not admit any AD-robust classifiers because the underlying classes are ambiguous. We provide criteria to determine whether a classification problem is well defined; describe the structure and attributes of AD-robust classifiers; introduce and explore the notions of conceptual entropy and regions of conceptual ambiguity for classifiers that are vulnerable to AD attacks; and discuss methods to bound the AD fooling rate of an attack. We define the notion of classifiers that exhibit hypersensitive behavior, that is, classifiers whose only mistakes are adversarial Doppelgängers. Improving the AD robustness of hypersensitive classifiers is equivalent to improving accuracy. We identify conditions guaranteeing that all classifiers with sufficiently high accuracy are hypersensitive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 133 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- Adversarial Robustness of Supervised Sparse CodingJeremias Sulam, Ramchandran Muthukumar, Raman AroraNeurIPS 2020 · 27 citations
Related papers
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang et al.ICLR 2024
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot et al.ICML 2020 · 103 citations
- Do Perceptually Aligned Gradients Imply Robustness?Roy Ganz, Bahjat Kawar, Michael EladICML 2023 · 18 citations
- Strong and Precise Modulation of Human Percepts via Robustified ANNsGuy Gaziv, Michael J. Lee, James J. DiCarloNeurIPS 2023 · 12 citations
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 61 citations
