Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient Descent
Oliver Bryniarski, Nabeel Hingun, Pedro Pachuca, Vincent Wang, Nicholas Carlini
摘要
Evading adversarial example detection defenses requires finding adversarial examples that must simultaneously (a) be misclassified by the model and (b) be detected as non-adversarial. We find that existing attacks that attempt to satisfy multiple simultaneous constraints often over-optimize against one constraint at the cost of satisfying another. We introduce Orthogonal Projected Gradient Descent, an improved attack technique to generate adversarial examples that avoids this problem by orthogonalizing the gradients when running standard gradient-based attacks. We use our technique to evade four state-of-the-art detection defenses, reducing their accuracy to 0% while maintaining a 0% detection rate. * Equal contributions. Authored alphabetically. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Increasing Confidence in Adversarial Robustness EvaluationsRoland S. Zimmermann, Wieland Brendel, Florian Tramèr, Nicholas CarliniNeurIPS 2022 · 被引用 25 次
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu 等ICML 2024 · 被引用 11 次
- Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsShawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng 等CCS 2022 · 被引用 9 次
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu 等ICLR 2026 · 被引用 3 次
- On the Robustness of Distributed Machine Learning Against Transfer AttacksSébastien Andreina, Pascal Zimmer, Ghassan KarameAAAI 2025 · 被引用 1 次
它引用的顶会 Paper8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
相关 Paper
- Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural NetworksWeiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter 等ICML 2022 · 被引用 5 次
- Guided Adversarial Attack for Evaluating and Enhancing Adversarial DefensesGaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, Venkatesh Babu R.NeurIPS 2020 · 被引用 123 次
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial ExamplesMaura Pintor, Luca Demetrio, Angelo Sotgiu, Ambra Demontis 等NeurIPS 2022 · 被引用 39 次
- Robust Adversarial Attacks Against Unknown Disturbance via Inverse Gradient SampleZhaoyang Zhang, Shen Wang, Runze Liu, Guopu Zhu 等ICLR 2026
- Detecting Adversarial Data Using Perturbation ForgeryQian Wang, Chen Li, Yuchen Luo, Hefei Ling 等CVPR 2025
