Stratified Adversarial Robustness with Rejection
Jiefeng Chen, Jayaram Raghuram, Jihye Choi, Xi Wu, Yingyu Liang, Somesh Jha
Abstract
Recently, there is an emerging interest in adversarially training a classifier with a rejection option (also known as a selective classifier) for boosting adversarial robustness. While rejection can incur a cost in many applications, existing studies typically associate zero cost with rejecting perturbed inputs, which can result in the rejection of numerous slightly-perturbed inputs that could be correctly classified. In this work, we study adversarially-robust classification with rejection in the stratified rejection setting, where the rejection cost is modeled by rejection loss functions monotonically non-increasing in the perturbation magnitude. We theoretically analyze the stratified rejection setting and propose a novel defense method -- Adversarial Training with Consistent Prediction-based Rejection (CPR) -- for building a robust selective classifier. Experiments on image datasets demonstrate that the proposed method significantly outperforms existing methods under strong adaptive attacks. For instance, on CIFAR-10, CPR reduces the total robust loss (for different rejection losses) by at least 7.3% under both seen and unseen attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3b12765-e94c-4b2c-adba-ae591dc5efeaCited by top-tier papers1
Ask how each one uses itBuilds on14
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Bag of Tricks for Adversarial TrainingTianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su et al.ICLR 2021 · 298 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
Related papers
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong et al.CVPR 2022 · 13 citations
- Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and RejectionNils Palumbo, Yang Guo, Xi Wu, Jiefeng Chen et al.ICML 2024
- Generalizing Consistent Multi-Class Classification with Rejection to be Compatible with Arbitrary LossesYuzhou Cao, Tianchi Cai, Lei Feng, Lihong Gu et al.NeurIPS 2022 · 42 citations
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 9 citations
- Randomization matters How to defend against strong adversarial attacksRafael Pinot, Raphael Ettedgui, Geovani Rizk, Yann Chevaleyre et al.ICML 2020 · 66 citations
