Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
Florian Tramèr
Abstract
Making classifiers robust to adversarial examples is challenging. Thus, many works tackle the seemingly easier task of detecting perturbed inputs. We show a barrier towards this goal. We prove a hardness reduction between detection and classification of adversarial examples: given a robust detector for attacks at distance (cid:15) (in some met-ric), we show how to build a similarly robust (but computationally inefficient) classifier for attacks at distance (cid:15)/ 2 . Our reduction is computationally inefficient , but preserves the sample complexity of the original detector. The reduction thus cannot be directly used to build practical classifiers. Instead, it is a useful sanity check to test whether empirical detection results imply something much stronger than the authors presumably anticipated (namely a highly robust and data-efficient classifier ). To illustrate, we revisit 14 empirical detector defenses published over the past years. For 12 / 14 defenses, we show that the claimed detection results imply an inefficient classifier with robustness far beyond the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong et al.CVPR 2022 · 13 citations
- Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust ClassifiersJonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso et al.NeurIPS 2024 · 12 citations
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi et al.CVPR 2026 · 12 citations
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu et al.ICML 2024 · 11 citations
Builds on10
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee et al.NDSS 2019 · 283 citations
Related papers
- Evaluating the Adversarial Robustness of Adaptive Test-time DefensesFrancesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer et al.ICML 2022 · 85 citations
- Computational Asymmetries in Robust ClassificationSamuele Marro, Michele LombardiICML 2023 · 2 citations
- Understanding the Limitations of Conditional Generative ModelsEthan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, Richard S. ZemelICLR 2020 · 65 citations
- Detection as Regression: Certified Object Detection with Median SmoothingPing-yeh Chiang, Michael J. Curry, Ahmed Abdelkader, Aounon Kumar et al.NeurIPS 2020 · 15 citations
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
