Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
David Stutz, Matthias Hein, Bernt Schiele
摘要
Adversarial training yields robust models against a specific threat model, e.g., adversarial examples. Typically robustness does not generalize to previously unseen threat models, e.g., other norms, or larger perturbations. Our confidence-calibrated adversarial training (CCAT) tackles this problem by biasing the model towards low confidence predictions on adversarial examples. By allowing to reject examples with low confidence, robustness generalizes beyond the threat model employed during training. CCAT, trained only on adversarial examples, increases robustness against larger , , and attacks, adversarial frames, distal adversarial examples and corrupted examples and yields better clean accuracy compared to adversarial training. For thorough evaluation we developed novel white- and black-box attacks directly attacking CCAT by maximizing confidence. For each threat model, we use attacks with up to restarts and iterations and report worst-case robust test error, extended to our confidence-thresholded setting, across all attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin 等ICML 2023 · 被引用 300 次
- Bag of Tricks for Adversarial TrainingTianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su 等ICLR 2021 · 被引用 298 次
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 被引用 217 次
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu 等ICML 2022 · 被引用 168 次
- Graph Posterior Network: Bayesian Predictive Uncertainty for Node ClassificationMaximilian Stadler, Bertrand Charpentier, Simon Geisler, Daniel Zügner 等NeurIPS 2021 · 被引用 133 次
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Sparse and Imperceivable Adversarial AttacksFrancesco Croce, Matthias HeinICCV 2019 · 被引用 228 次
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson 等AAAI 2020 · 被引用 210 次
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 被引用 171 次
相关 Paper
- Adversarial Robustness against Multiple and Single lp-Threat Models via Quick Fine-Tuning of Robust ClassifiersFrancesco Croce, Matthias HeinICML 2022 · 被引用 26 次
- Seasoning Model Soups for Robustness to Adversarial and Natural Distribution ShiftsFrancesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, Sven GowalCVPR 2023
- Inequality phenomenon in l∞-adversarial training, and its unrealized threatsRanjie Duan, Yuefeng Chen, Yao Zhu, Xiaojun Jia 等ICLR 2023
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong 等CVPR 2022 · 被引用 13 次
- Sample Efficient Detection and Classification of Adversarial Attacks via Self-Supervised EmbeddingsMazda Moayeri, Soheil FeiziICCV 2021 · 被引用 20 次
