Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural Networks
Zhen Lin, Shubhendu Trivedi, Jimeng Sun
Abstract
Deep neural network (DNN) classifiers are often overconfident, producing miscalibrated class probabilities. In high-risk applications like healthcare, practitioners require fully calibrated probability predictions for decision-making. That is, conditioned on the prediction vector, every class' probability should be close to the predicted value. Most existing calibration methods either lack theoretical guarantees for producing calibrated outputs, reduce classification accuracy in the process, or only calibrate the predicted class. This paper proposes a new Kernel-based calibration method called KCal. Unlike existing calibration procedures, KCal does not operate directly on the logits or softmax outputs of the DNN. Instead, KCal learns a metric space on the penultimate-layer latent embedding and generates predictions using kernel density estimates on a calibration set. We first analyze KCal theoretically, showing that it enjoys a provable full calibration guarantee. Then, through extensive experiments across a variety of datasets, we show that KCal consistently outperforms baselines as measured by the calibration error and by proper scoring rules like the Brier Score. Recent research effort has started to focus on full calibration, for example, in Vaicenavicius et al. (2019); Kull et al. (2019); Widmann et al. (2019); Karandikar et al. (2021); Mukhoti et al. (2020); Patel et al. (2021). We approach this problem by leveraging the latent neural network embedding in a nonparametric manner. Nonparametric methods such as histogram binning (HB) (Zadrozny & Elkan, 2001) and isotonic regression (IR) (Zadrozny & Elkan, 2002), are natural for calibration and have become popular. Gupta & Ramdas (2021) recently showed a calibration guarantee for HB. However, HB usually leads to noticeable drops in accuracy (Patel et al., 2021), and IR is prone to overfitting (Niculescu-Mizil & Caruana, 2005) . Unlike existing methods, we take one step back and train a new low-dimensional metric space on the penultimate-layer embeddings of DNNs. Then, we use a kernel density estimationbased classifier to predict the class probabilities directly. We refer to our Kernel-based Calibration method as KCal. Unlike most calibration methods, KCal provides high probability error bounds for full calibration under standard assumptions. Empirically, we show that with little overhead, KCal outperforms all existing calibration methods in terms of calibration quality, across multiple tasks and DNN architectures, while maintaining and sometimes improving the classification accuracy. Summary of Contributions: • We propose KCal, a principled method that calibrates DNNs using kernel density estimation on the latent embeddings. • We present an efficient pipeline to train KCal, including a dimension-reducing projection and a stratified sampling method to facilitate efficient training. • We provide finite sample bounds for the calibration error of KCal-calibrated output under standard assumptions. To the best of our knowledge, this is the first method with a full calibration guarantee. • In extensive experiments on multiple datasets and state-of-the-art models, we found that KCal outperforms existing calibration methods in commonly used evaluation metrics. We also show that KCal provides more reliable predictions for important classes in the healthcare datasets. The code to replicate all our experimental results is submitted along with supplementary materials. RELATED WORK Research on calibration originated in the context of meteorology and weather forecasting (see Murphy & Winkler (1984) for an overview) and has a long history, much older than the field of machine
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bae98f9-ec6e-41c9-8ced-85924f9d75e3Cited by top-tier papers4
- Proximity-Informed Calibration for Deep Neural NetworksMiao Xiong, Ailin Deng, Pang Wei Koh, Jiaying Wu et al.NeurIPS 2023 · 30 citations
- Confidence Calibration of Classifiers with Many ClassesAdrien Le-Coz, Stéphane Herbin, Faouzi AdjedNeurIPS 2024 · 21 citations
- Expert load matters: operating networks at high accuracy and low manual effortSara Sangalli, Ertunc Erdil, Ender KonukogluNeurIPS 2023 · 10 citations
- PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep LearningJohn Wu, Yongda Fan, Zhenbang Wu, Paul Landes et al.ICML 2026 · 1 citation
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
- Calibration of Neural Networks using SplinesKartik Gupta, Amir Rahimi, Thalaiyasingam Ajanthan, Thomas Mensink et al.ICLR 2021 · 128 citations
Related papers
- A Consistent and Differentiable Lp Canonical Calibration Error EstimatorTeodora Popordanoska, Raphael Sayer, Matthew B. BlaschkoNeurIPS 2022 · 58 citations
- Calibrated Reliable Regression using Maximum Mean DiscrepancyPeng Cui, Wenbo Hu, Jun ZhuNeurIPS 2020 · 71 citations
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width LimitBen Adlam, Jaehoon Lee, Lechao Xiao, Jeffrey Pennington et al.ICLR 2021 · 3 citations
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
- Distance-Based Learning from Errors for Confidence CalibrationChen Xing, Sercan Ömer Arik, Zizhao Zhang, Tomas PfisterICLR 2020 · 42 citations
