Doppelgangers and Adversarial Vulnerability
George I. Kamberov
摘要
Machine learning (ML) classifiers can make mistakes that are perceptually and cognitively disturbing to humans. The most notorious examples of such errors are adversarial visual metamers. This paper investigates the phenomenon of adversarial Doppelgängers (AD), which encompasses adversarial visual metamers, and compares the performance and robustness of ML classifiers to human performance. We find that ADs are inputs that are close to each other with respect to a perceptual metric defined in this paper, and show that ADs are qualitatively different from the usual adversarial examples. The vast majority of classifiers are vulnerable to ADs and robustness-accuracy trade-offs may not improve them. Some classification problems do not admit any AD-robust classifiers because the underlying classes are ambiguous. We provide criteria to determine whether a classification problem is well defined; describe the structure and attributes of AD-robust classifiers; introduce and explore the notions of conceptual entropy and regions of conceptual ambiguity for classifiers that are vulnerable to AD attacks; and discuss methods to bound the AD fooling rate of an attack. We define the notion of classifiers that exhibit hypersensitive behavior, that is, classifiers whose only mistakes are adversarial Doppelgängers. Improving the AD robustness of hypersensitive classifiers is equivalent to improving accuracy. We identify conditions guaranteeing that all classifiers with sufficiently high accuracy are hypersensitive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao 等ICML 2022 · 被引用 663 次
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 被引用 133 次
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 被引用 83 次
- Adversarial Robustness of Supervised Sparse CodingJeremias Sulam, Ramchandran Muthukumar, Raman AroraNeurIPS 2020 · 被引用 27 次
相关 Paper
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
- Do Perceptually Aligned Gradients Imply Robustness?Roy Ganz, Bahjat Kawar, Michael EladICML 2023 · 被引用 18 次
- Strong and Precise Modulation of Human Percepts via Robustified ANNsGuy Gaziv, Michael J. Lee, James J. DiCarloNeurIPS 2023 · 被引用 12 次
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 被引用 61 次
