The Devil is in the Margin: Margin-based Label Smoothing for Network Calibration
Bingyuan Liu, Ismail Ben Ayed, Adrian Galdran, Jose Dolz
摘要
In spite of the dominant performances of deep neural networks, recent works have shown that they are poorly calibrated, resulting in over-confident predictions. Miscalibration can be exacerbated by overfitting due to the minimization of the cross-entropy during training, as it promotes the predicted softmax probabilities to match the one-hot label assignments. This yields a pre-softmax activation of the correct class that is significantly larger than the remaining activations. Recent evidence from the literature suggests that loss functions that embed implicit or explicit maximization of the entropy of predictions yield state-of-the-art calibration performances. We provide a unifying constrained-optimization perspective of current state-of-the-art calibration losses. Specifically, these losses could be viewed as approximations of a linear penalty (or a Lagrangian term) imposing equality constraints on logit distances. This points to an important limitation of such underlying equality constraints, whose ensuing gradients constantly push towards a non-informative solution, which might prevent from reaching the best compromise between the discriminative performance and calibration of the model during gradient-based optimization. Following our observations, we propose a simple and flexible generalization based on inequality constraints, which imposes a controllable margin on logit distances. Comprehensive experiments on a variety of image classification, semantic segmentation and NLP benchmarks demonstrate that our method sets novel state-of-the-art results on these tasks in terms of network calibration, without affecting the discriminative performance. The code is available at https://github.com/by-liu/MbLS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Cal-DETR: Calibrated Detection TransformerMuhammad Akhtar Munir, Salman H. Khan, Muhammad Haris Khan, Mohsen Ali 等NeurIPS 2023 · 被引用 24 次
- ACLS: Adaptive and Conditional Label Smoothing for Network CalibrationHyekang Park, Jongyoun Noh, Youngmin Oh, Donghyeon Baek 等ICCV 2023 · 被引用 22 次
- RankMixup: Ranking-Based Mixup Training for Network CalibrationJongyoun Noh, Hyekang Park, Junghyup Lee, Bumsub HamICCV 2023 · 被引用 22 次
- LFME: A Simple Framework for Learning from Multiple Experts in Domain GeneralizationLiang Chen, Yong Zhang, Yibing Song, Zhiqiang Shen 等NeurIPS 2024 · 被引用 15 次
- Adaptive Decision Boundary for Few-Shot Class-Incremental LearningLinhao Li, Yongzhang Tan, Siyuan Yang, Hao Cheng 等AAAI 2025 · 被引用 10 次
它引用的顶会 Paper7
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
- Local Temperature Scaling for Probability CalibrationZhipeng Ding, Xu Han, Peirong Liu, Marc NiethammerICCV 2021 · 被引用 109 次
相关 Paper
- Class Adaptive Network CalibrationBingyuan Liu, Jérôme Rony, Adrian Galdran, Jose Dolz 等CVPR 2023
- Calibrating Deep Neural Networks by Pairwise ConstraintsJiacheng Cheng, Nuno VasconcelosCVPR 2022 · 被引用 21 次
- Dual Focal Loss for CalibrationLinwei Tao, Minjing Dong, Chang XuICML 2023 · 被引用 56 次
- AdaFocal: Calibration-aware Adaptive Focal LossArindam Ghosh, Thomas Schaaf, Matthew GormleyNeurIPS 2022 · 被引用 71 次
- Uncertainty Weighted Gradients for Model CalibrationJinxu Lin, Linwei Tao, Minjing Dong, Chang XuCVPR 2025
