Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient Descent
Oliver Bryniarski, Nabeel Hingun, Pedro Pachuca, Vincent Wang, Nicholas Carlini
Abstract
Evading adversarial example detection defenses requires finding adversarial examples that must simultaneously (a) be misclassified by the model and (b) be detected as non-adversarial. We find that existing attacks that attempt to satisfy multiple simultaneous constraints often over-optimize against one constraint at the cost of satisfying another. We introduce Orthogonal Projected Gradient Descent, an improved attack technique to generate adversarial examples that avoids this problem by orthogonalizing the gradients when running standard gradient-based attacks. We use our technique to evade four state-of-the-art detection defenses, reducing their accuracy to 0% while maintaining a 0% detection rate. * Equal contributions. Authored alphabetically. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adde4485-ca8c-4622-a7c9-bbc5042b196dCited by top-tier papers9
- Increasing Confidence in Adversarial Robustness EvaluationsRoland S. Zimmermann, Wieland Brendel, Florian Tramèr, Nicholas CarliniNeurIPS 2022 · 25 citations
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu et al.ICML 2024 · 11 citations
- Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsShawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng et al.CCS 2022 · 9 citations
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu et al.ICLR 2026 · 3 citations
- On the Robustness of Distributed Machine Learning Against Transfer AttacksSébastien Andreina, Pascal Zimmer, Ghassan KarameAAAI 2025 · 1 citation
Builds on8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
Related papers
- Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural NetworksWeiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter et al.ICML 2022 · 5 citations
- Guided Adversarial Attack for Evaluating and Enhancing Adversarial DefensesGaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, Venkatesh Babu R.NeurIPS 2020 · 123 citations
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial ExamplesMaura Pintor, Luca Demetrio, Angelo Sotgiu, Ambra Demontis et al.NeurIPS 2022 · 39 citations
- Robust Adversarial Attacks Against Unknown Disturbance via Inverse Gradient SampleZhaoyang Zhang, Shen Wang, Runze Liu, Guopu Zhu et al.ICLR 2026
- Detecting Adversarial Data Using Perturbation ForgeryQian Wang, Chen Li, Yuchen Luo, Hefei Ling et al.CVPR 2025
