Towards Certification of Uncertainty Calibration under Adversarial Attacks
Cornelius Emde, Francesco Pinto, Thomas Lukasiewicz, Philip Torr, Adel Bibi
Abstract
Since neural classifiers are known to be sensitive to adversarial perturbations that alter their accuracy, certification methods have been developed to provide provable guarantees on the insensitivity of their predictions to such perturbations. Furthermore, in safety-critical applications, the frequentist interpretation of the confidence of a classifier (also known as model calibration) can be of utmost importance. This property can be measured via the Brier score or the expected calibration error. We show that attacks can significantly harm calibration, and thus propose certified calibration as worst-case bounds on calibration under adversarial perturbations. Specifically, we produce analytic bounds for the Brier score and approximate bounds via the solution of a mixed-integer program on the expected calibration error. Finally, we propose novel calibration attacks and demonstrate how they can improve model calibration through adversarial calibration training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ff85a28-4fc0-4856-b4d7-dcad0a3dceafBuilds on15
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- MACER: Attack-free and Scalable Robust Training via Maximizing Certified RadiusRuntian Zhai, Chen Dan, Di He, Huan Zhang et al.ICLR 2020 · 195 citations
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 186 citations
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 158 citations
- Soft Calibration Objectives for Neural NetworksArchit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan et al.NeurIPS 2021 · 127 citations
Related papers
- Support is All You Need for Certified VAE TrainingChangming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep SinghICLR 2025
- Certifiably Adversarially Robust Detection of Out-of-Distribution DataJulian Bitterwolf, Alexander Meinke, Matthias HeinNeurIPS 2020 · 91 citations
- Regularized Training and Tight Certification for Randomized Smoothed Classifier with Provable RobustnessHuijie Feng, Chunpeng Wu, Guoyang Chen, Weifeng Zhang et al.AAAI 2020 · 13 citations
- Probabilistic Robustness Certificates against Adversarial AttacksSara Taheri, Majid ZamaniICML 2026
- The Confidence Trap: Calibration Attacks for Graph Neural NetworksCuong Dang, Jiahao Zhang, Hieu Ta Quang, Dung Le et al.KDD 2026
