On the Limitations of Temperature Scaling for Distributions with Overlaps
Muthu Chidambaram, Rong Ge
Abstract
Despite the impressive generalization capabilities of deep neural networks, they have been repeatedly shown to be overconfident when they are wrong. Fixing this issue is known as model calibration, and has consequently received much attention in the form of modified training schemes and post-training calibration procedures such as temperature scaling. While temperature scaling is frequently used because of its simplicity, it is often outperformed by modified training schemes. In this work, we identify a specific bottleneck for the performance of temperature scaling. We show that for empirical risk minimizers for a general set of distributions in which the supports of classes have overlaps, the performance of temperature scaling degrades with the amount of overlap between classes, and asymptotically becomes no better than random when there are a large number of classes. On the other hand, we prove that optimizing a modified form of the empirical risk induced by the Mixup data augmentation technique can in fact lead to reasonably good calibration performance, showing that training-time calibration may be necessary in some situations. We also verify that our theoretical results reflect practice by showing that Mixup significantly outperforms empirical risk minimization (with respect to multiple calibration metrics) on image classification benchmarks with class overlaps introduced in the form of label noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen et al.AAAI 2026
- T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental LearningSeonghyeon Hwang, Minsu Kim, Steven Euijong WhangCVPR 2025
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
Builds on9
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
- Local Temperature Scaling for Probability CalibrationZhipeng Ding, Xu Han, Peirong Liu, Marc NiethammerICCV 2021 · 109 citations
Related papers
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 250 citations
- Rethinking Data Distillation: Do Not Overlook CalibrationDongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie et al.ICCV 2023 · 19 citations
- On the Pitfall of Mixup for Uncertainty CalibrationDeng-Bao Wang, Lanqing Li, Peilin Zhao, Pheng-Ann Heng et al.CVPR 2023
- Tailoring Mixup to Data for CalibrationQuentin Bouniot, Pavlo Mozharovskyi, Florence d'Alché-BucICLR 2025
- Set Learning for Accurate and Calibrated ModelsLukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang, Thomas Unterthiner et al.ICLR 2024 · 4 citations
