Provably robust classification of adversarial examples with detection
Fatemeh Sheikholeslami, Ali Lotfi, J. Zico Kolter
摘要
Adversarial attacks against deep networks can be defended against either by building robust classifiers or, by creating classifiers that can detect the presence of adversarial perturbations. Although it may intuitively seem easier to simply detect attacks rather than build a robust classifier, this has not bourne out in practice even empirically, as most detection methods have subsequently been broken by adaptive attacks, thus necessitating verifiable performance for detection mechanisms. In this paper, we propose a new method for jointly training a provably robust classifier and detector. Specifically, we show that by introducing an additional "abstain/detection" into a classifier, we can modify existing certified defense mechanisms to allow the classifier to either robustly classify or detect adversarial attacks. We extend the common interval bound propagation (IBP) method for certified robustness under perturbations to account for our new robust objective, and show that the method outperforms traditional IBP used in isolation, especially for large perturbation sizes. Specifically, tests on MNIST and CIFAR-10 datasets exhibit promising results, for example with provable robust error less than and , for and natural error, for and on the CIFAR-10 dataset, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying ThemFlorian TramèrICML 2022 · 被引用 82 次
- Scalable Certified Segmentation via Randomized SmoothingMarc Fischer, Maximilian Baader, Martin T. VechevICML 2021 · 被引用 49 次
- Adversarial Attack Generation Empowered by Min-Max OptimizationJingkang Wang, Tianyun Zhang, Sijia Liu, Pin-Yu Chen 等NeurIPS 2021 · 被引用 49 次
- Synergy-of-Experts: Collaborate to Improve Adversarial RobustnessSen Cui, Jingfeng Zhang, Jian Liang, Bo Han 等NeurIPS 2022 · 被引用 12 次
- Stratified Adversarial Robustness with RejectionJiefeng Chen, Jayaram Raghuram, Jihye Choi, Xi Wu 等ICML 2023 · 被引用 4 次
相关 Paper
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel 等ICCV 2019 · 被引用 196 次
- Certifiably Adversarially Robust Detection of Out-of-Distribution DataJulian Bitterwolf, Alexander Meinke, Matthias HeinNeurIPS 2020 · 被引用 91 次
- Towards Better Understanding of Training Certifiably Robust Models against Adversarial ExamplesSungyoon Lee, Woojin Lee, Jinseong Park, Jaewook LeeNeurIPS 2021 · 被引用 27 次
- On the Convergence of Certified Robust Training with Interval Bound PropagationYihan Wang, Zhouxing Shi, Quanquan Gu, Cho-Jui HsiehICLR 2022 · 被引用 11 次
