Robust Models are less Over-Confident
Julia Grabinski, Paul Gavrikov, Janis Keuper, Margret Keuper
Abstract
Despite the success of convolutional neural networks (CNNs) in many academic benchmarks for computer vision tasks, their application in the real-world is still facing fundamental challenges. One of these open problems is the inherent lack of robustness, unveiled by the striking effectiveness of adversarial attacks. Current attack methods are able to manipulate the network's prediction by adding specific but small amounts of noise to the input. In turn, adversarial training (AT) aims to achieve robustness against such attacks and ideally a better model generalization ability by including adversarial samples in the trainingset. However, an in-depth analysis of the resulting robust models beyond adversarial robustness is still pending. In this paper, we empirically analyze a variety of adversarially trained models that achieve high robust accuracies when facing state-of-the-art attacks and we show that AT has an interesting side-effect: it leads to models that are significantly less overconfident with their decisions, even on clean data than non-robust models. Further, our analysis of robust models shows that not only AT but also the model's building blocks (like activation functions and pooling) have a strong influence on the models' prediction confidences. Data&Project website: https://github.com/GeJulia/robustness_confidences_evaluation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc22ce9d-9e9b-4352-a245-be5294523157Cited by top-tier papers6
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language ModelsYihao Zhang, Zeming Wei, Jun Sun, Meng SunNeurIPS 2024 · 16 citations
- Annealing Self-Distillation Rectification Improves Adversarial TrainingYu-Yu Wu, Hung-Jui Wang, Shang-Tse ChenICLR 2024 · 10 citations
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable RewardsZhengzhao Ma, Xueru Wen, Boxi Cao, Yaojie Lu et al.ICML 2026 · 5 citations
- Towards Certification of Uncertainty Calibration under Adversarial AttacksCornelius Emde, Francesco Pinto, Thomas Lukasiewicz, Philip Torr et al.ICLR 2025 · 1 citation
- Revisiting Unknowns: Towards Effective and Efficient Open-Set Active LearningChen-Chen Zong, Yu-Qi Chi, Xie-Yang Wang, Yan Cui et al.CVPR 2026 · 1 citation
Builds on26
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Adversarial Training on Purification (AToP): Advancing Both Robustness and GeneralizationGuang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka et al.ICLR 2024 · 25 citations
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 158 citations
- On the Duality Between Sharpness-Aware Minimization and Adversarial TrainingYihao Zhang, Hangzhou He, Jingyu Zhu, Huanran Chen et al.ICML 2024 · 29 citations
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 3 citations
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
