Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
Florian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot, Jörn-Henrik Jacobsen
摘要
Adversarial examples are malicious inputs crafted to induce misclassification. Commonly studied sensitivity-based adversarial examples introduce semantically-small changes to an input that result in a different model prediction. This paper studies a complementary failure mode, invariance-based adversarial examples, that introduce minimal semantic changes that modify an input's true label yet preserve the model's prediction. We demonstrate fundamental tradeoffs between these two types of adversarial examples. We show that defenses against sensitivity-based attacks actively harm a model's accuracy on invariance-based attacks, and that new approaches are needed to resist both attack types. In particular, we break state-of-the-art adversarially-trained and certifiably-robust models by generating small perturbations that the models are (provably) robust to, yet that change an input's class according to human labelers. Finally, we formally show that the existence of excessively invariant classifiers arises from the presence of overly-robust predictive features in standard datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg 等NeurIPS 2021 · 被引用 384 次
- Blind Backdoors in Deep Learning ModelsEugene Bagdasaryan, Vitaly ShmatikovUSENIX Security 2021 · 被引用 372 次
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer 等NeurIPS 2021 · 被引用 304 次
- To be Robust or to be Fair: Towards Fairness in Adversarial TrainingHan Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain 等ICML 2021 · 被引用 218 次
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu 等ICML 2022 · 被引用 168 次
它引用的顶会 Paper3
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional NetworksKenneth T. Co, Luis Muñoz-González, Sixte de Maupeou, Emil C. LupuCCS 2019 · 被引用 77 次
- AdVersarial: Perceptual Ad Blocking meets Adversarial Machine LearningFlorian Tramèr, Pascal Dupré, Gili Rusak, Giancarlo Pellegrino 等CCS 2019 · 被引用 65 次
相关 Paper
- Balanced Adversarial Training: Balancing Tradeoffs between Fickleness and Obstinacy in NLP ModelsHannah Chen, Yangfeng Ji, David E. EvansEMNLP 2022 · 被引用 4 次
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 被引用 61 次
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
- Breaking Certified Defenses: Semantic Adversarial Examples with Spoofed robustness CertificatesAmin Ghiasi, Ali Shafahi, Tom GoldsteinICLR 2020 · 被引用 57 次
