Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
David Stutz, Matthias Hein, Bernt Schiele
Abstract
Adversarial training yields robust models against a specific threat model, e.g., adversarial examples. Typically robustness does not generalize to previously unseen threat models, e.g., other norms, or larger perturbations. Our confidence-calibrated adversarial training (CCAT) tackles this problem by biasing the model towards low confidence predictions on adversarial examples. By allowing to reject examples with low confidence, robustness generalizes beyond the threat model employed during training. CCAT, trained only on adversarial examples, increases robustness against larger , , and attacks, adversarial frames, distal adversarial examples and corrupted examples and yields better clean accuracy compared to adversarial training. For thorough evaluation we developed novel white- and black-box attacks directly attacking CCAT by maximizing confidence. For each threat model, we use attacks with up to restarts and iterations and report worst-case robust test error, extended to our confidence-thresholded setting, across all attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2e07179-bee0-457e-acfd-3b40ea71eb6cCited by top-tier papers39
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- Bag of Tricks for Adversarial TrainingTianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su et al.ICLR 2021 · 298 citations
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 217 citations
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu et al.ICML 2022 · 168 citations
- Graph Posterior Network: Bayesian Predictive Uncertainty for Node ClassificationMaximilian Stadler, Bertrand Charpentier, Simon Geisler, Daniel Zügner et al.NeurIPS 2021 · 133 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Sparse and Imperceivable Adversarial AttacksFrancesco Croce, Matthias HeinICCV 2019 · 228 citations
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson et al.AAAI 2020 · 210 citations
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 171 citations
Related papers
- Adversarial Robustness against Multiple and Single lp-Threat Models via Quick Fine-Tuning of Robust ClassifiersFrancesco Croce, Matthias HeinICML 2022 · 26 citations
- Seasoning Model Soups for Robustness to Adversarial and Natural Distribution ShiftsFrancesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, Sven GowalCVPR 2023
- Inequality phenomenon in l∞-adversarial training, and its unrealized threatsRanjie Duan, Yuefeng Chen, Yao Zhu, Xiaojun Jia et al.ICLR 2023
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong et al.CVPR 2022 · 13 citations
- Sample Efficient Detection and Classification of Adversarial Attacks via Self-Supervised EmbeddingsMazda Moayeri, Soheil FeiziICCV 2021 · 20 citations
