Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
Florian Tramèr
摘要
Making classifiers robust to adversarial examples is challenging. Thus, many works tackle the seemingly easier task of detecting perturbed inputs. We show a barrier towards this goal. We prove a hardness reduction between detection and classification of adversarial examples: given a robust detector for attacks at distance (cid:15) (in some met-ric), we show how to build a similarly robust (but computationally inefficient) classifier for attacks at distance (cid:15)/ 2 . Our reduction is computationally inefficient , but preserves the sample complexity of the original detector. The reduction thus cannot be directly used to build practical classifiers. Instead, it is a useful sanity check to test whether empirical detection results imply something much stronger than the authors presumably anticipated (namely a highly robust and data-efficient classifier ). To illustrate, we revisit 14 empirical detector defenses published over the past years. For 12 / 14 defenses, we show that the claimed detection results imply an inefficient classifier with robustness far beyond the state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong 等CVPR 2022 · 被引用 13 次
- Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust ClassifiersJonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso 等NeurIPS 2024 · 被引用 12 次
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi 等CVPR 2026 · 被引用 12 次
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu 等ICML 2024 · 被引用 11 次
它引用的顶会 Paper10
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee 等NDSS 2019 · 被引用 283 次
相关 Paper
- Evaluating the Adversarial Robustness of Adaptive Test-time DefensesFrancesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer 等ICML 2022 · 被引用 85 次
- Computational Asymmetries in Robust ClassificationSamuele Marro, Michele LombardiICML 2023 · 被引用 2 次
- Understanding the Limitations of Conditional Generative ModelsEthan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, Richard S. ZemelICLR 2020 · 被引用 65 次
- Detection as Regression: Certified Object Detection with Median SmoothingPing-yeh Chiang, Michael J. Curry, Ahmed Abdelkader, Aounon Kumar 等NeurIPS 2020 · 被引用 15 次
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
