On the Limitations of Temperature Scaling for Distributions with Overlaps
Muthu Chidambaram, Rong Ge
摘要
Despite the impressive generalization capabilities of deep neural networks, they have been repeatedly shown to be overconfident when they are wrong. Fixing this issue is known as model calibration, and has consequently received much attention in the form of modified training schemes and post-training calibration procedures such as temperature scaling. While temperature scaling is frequently used because of its simplicity, it is often outperformed by modified training schemes. In this work, we identify a specific bottleneck for the performance of temperature scaling. We show that for empirical risk minimizers for a general set of distributions in which the supports of classes have overlaps, the performance of temperature scaling degrades with the amount of overlap between classes, and asymptotically becomes no better than random when there are a large number of classes. On the other hand, we prove that optimizing a modified form of the empirical risk induced by the Mixup data augmentation technique can in fact lead to reasonably good calibration performance, showing that training-time calibration may be necessary in some situations. We also verify that our theoretical results reflect practice by showing that Mixup significantly outperforms empirical risk minimization (with respect to multiple calibration metrics) on image classification benchmarks with class overlaps introduced in the form of label noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen 等AAAI 2026
- T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental LearningSeonghyeon Hwang, Minsu Kim, Steven Euijong WhangCVPR 2025
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
它引用的顶会 Paper9
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 被引用 177 次
- Local Temperature Scaling for Probability CalibrationZhipeng Ding, Xu Han, Peirong Liu, Marc NiethammerICCV 2021 · 被引用 109 次
相关 Paper
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 被引用 250 次
- Rethinking Data Distillation: Do Not Overlook CalibrationDongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie 等ICCV 2023 · 被引用 19 次
- On the Pitfall of Mixup for Uncertainty CalibrationDeng-Bao Wang, Lanqing Li, Peilin Zhao, Pheng-Ann Heng 等CVPR 2023
- Tailoring Mixup to Data for CalibrationQuentin Bouniot, Pavlo Mozharovskyi, Florence d'Alché-BucICLR 2025
- Set Learning for Accurate and Calibrated ModelsLukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang, Thomas Unterthiner 等ICLR 2024 · 被引用 4 次
