The Devil is in the Margin: Margin-based Label Smoothing for Network Calibration
Bingyuan Liu, Ismail Ben Ayed, Adrian Galdran, Jose Dolz
Abstract
In spite of the dominant performances of deep neural networks, recent works have shown that they are poorly calibrated, resulting in over-confident predictions. Miscalibration can be exacerbated by overfitting due to the minimization of the cross-entropy during training, as it promotes the predicted softmax probabilities to match the one-hot label assignments. This yields a pre-softmax activation of the correct class that is significantly larger than the remaining activations. Recent evidence from the literature suggests that loss functions that embed implicit or explicit maximization of the entropy of predictions yield state-of-the-art calibration performances. We provide a unifying constrained-optimization perspective of current state-of-the-art calibration losses. Specifically, these losses could be viewed as approximations of a linear penalty (or a Lagrangian term) imposing equality constraints on logit distances. This points to an important limitation of such underlying equality constraints, whose ensuing gradients constantly push towards a non-informative solution, which might prevent from reaching the best compromise between the discriminative performance and calibration of the model during gradient-based optimization. Following our observations, we propose a simple and flexible generalization based on inequality constraints, which imposes a controllable margin on logit distances. Comprehensive experiments on a variety of image classification, semantic segmentation and NLP benchmarks demonstrate that our method sets novel state-of-the-art results on these tasks in terms of network calibration, without affecting the discriminative performance. The code is available at https://github.com/by-liu/MbLS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99f829f5-7ef7-41c8-919a-380468c606c7Cited by top-tier papers31
- Cal-DETR: Calibrated Detection TransformerMuhammad Akhtar Munir, Salman H. Khan, Muhammad Haris Khan, Mohsen Ali et al.NeurIPS 2023 · 24 citations
- ACLS: Adaptive and Conditional Label Smoothing for Network CalibrationHyekang Park, Jongyoun Noh, Youngmin Oh, Donghyeon Baek et al.ICCV 2023 · 22 citations
- RankMixup: Ranking-Based Mixup Training for Network CalibrationJongyoun Noh, Hyekang Park, Junghyup Lee, Bumsub HamICCV 2023 · 22 citations
- LFME: A Simple Framework for Learning from Multiple Experts in Domain GeneralizationLiang Chen, Yong Zhang, Yibing Song, Zhiqiang Shen et al.NeurIPS 2024 · 15 citations
- Adaptive Decision Boundary for Few-Shot Class-Incremental LearningLinhao Li, Yongzhang Tan, Siyuan Yang, Hao Cheng et al.AAAI 2025 · 10 citations
Builds on7
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
- Local Temperature Scaling for Probability CalibrationZhipeng Ding, Xu Han, Peirong Liu, Marc NiethammerICCV 2021 · 109 citations
Related papers
- Class Adaptive Network CalibrationBingyuan Liu, Jérôme Rony, Adrian Galdran, Jose Dolz et al.CVPR 2023
- Calibrating Deep Neural Networks by Pairwise ConstraintsJiacheng Cheng, Nuno VasconcelosCVPR 2022 · 21 citations
- Dual Focal Loss for CalibrationLinwei Tao, Minjing Dong, Chang XuICML 2023 · 56 citations
- AdaFocal: Calibration-aware Adaptive Focal LossArindam Ghosh, Thomas Schaaf, Matthew GormleyNeurIPS 2022 · 71 citations
- Uncertainty Weighted Gradients for Model CalibrationJinxu Lin, Linwei Tao, Minjing Dong, Chang XuCVPR 2025
