Robust Models are less Over-Confident
Julia Grabinski, Paul Gavrikov, Janis Keuper, Margret Keuper
摘要
Despite the success of convolutional neural networks (CNNs) in many academic benchmarks for computer vision tasks, their application in the real-world is still facing fundamental challenges. One of these open problems is the inherent lack of robustness, unveiled by the striking effectiveness of adversarial attacks. Current attack methods are able to manipulate the network's prediction by adding specific but small amounts of noise to the input. In turn, adversarial training (AT) aims to achieve robustness against such attacks and ideally a better model generalization ability by including adversarial samples in the trainingset. However, an in-depth analysis of the resulting robust models beyond adversarial robustness is still pending. In this paper, we empirically analyze a variety of adversarially trained models that achieve high robust accuracies when facing state-of-the-art attacks and we show that AT has an interesting side-effect: it leads to models that are significantly less overconfident with their decisions, even on clean data than non-robust models. Further, our analysis of robust models shows that not only AT but also the model's building blocks (like activation functions and pooling) have a strong influence on the models' prediction confidences. Data&Project website: https://github.com/GeJulia/robustness_confidences_evaluation
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language ModelsYihao Zhang, Zeming Wei, Jun Sun, Meng SunNeurIPS 2024 · 被引用 16 次
- Annealing Self-Distillation Rectification Improves Adversarial TrainingYu-Yu Wu, Hung-Jui Wang, Shang-Tse ChenICLR 2024 · 被引用 10 次
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable RewardsZhengzhao Ma, Xueru Wen, Boxi Cao, Yaojie Lu 等ICML 2026 · 被引用 5 次
- Towards Certification of Uncertainty Calibration under Adversarial AttacksCornelius Emde, Francesco Pinto, Thomas Lukasiewicz, Philip Torr 等ICLR 2025 · 被引用 1 次
- Revisiting Unknowns: Towards Effective and Efficient Open-Set Active LearningChen-Chen Zong, Yu-Qi Chi, Xie-Yang Wang, Yan Cui 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper26
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
相关 Paper
- Adversarial Training on Purification (AToP): Advancing Both Robustness and GeneralizationGuang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka 等ICLR 2024 · 被引用 25 次
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 被引用 158 次
- On the Duality Between Sharpness-Aware Minimization and Adversarial TrainingYihao Zhang, Hangzhou He, Jingyu Zhu, Huanran Chen 等ICML 2024 · 被引用 29 次
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 被引用 3 次
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 被引用 1 次
