Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It
Guoxuan Xia, Olivier Laurent, Gianni Franchi, Christos-Savvas Bouganis
Abstract
Label smoothing (LS) is a popular regularisation method for training neural networks as it is effective in improving test accuracy and is simple to implement. "Hard" one-hot labels are "smoothed" by uniformly distributing probability mass to other classes, reducing overfitting. Prior work has suggested that in some cases LS can degrade selective classification (SC) -where the aim is to reject misclassifications using a model's uncertainty. In this work, we first demonstrate empirically across an extended range of large-scale tasks and architectures that LS consistently degrades SC. We then address a gap in existing knowledge, providing an explanation for this behaviour by analysing logit-level gradients: LS degrades the uncertainty rank ordering of correct vs incorrect predictions by suppressing the max logit more when a prediction is likely to be correct, and less when it is likely to be wrong. This elucidates previously reported experimental results where strong classifiers underperform in SC. We then demonstrate the empirical effectiveness of post-hoc logit normalisation for recovering lost SC performance caused by LS. Furthermore, linking back to our gradient analysis, we again provide an explanation for why such normalisation is effective. Project page: https://ensta-u2is-ai.github. io/Understanding-Label-smoothing-Selective-classification/ Recently Zhu et al. (2022; 2024) empirically observe that LS can lead to worse SC for convolutional neural networks (CNNs) performing image classification, yet it is not clear why this occurs. Concurrently, large-scale empirical evaluations of open-source pre-trained models have shown
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab709b77-ced0-4715-93da-e7f6e3c3ac3cCited by top-tier papers2
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
- Towards Understanding and Quantifying Uncertainty for Text-to-Image GenerationGianni Franchi, Nacim Belkhir, Dat Nguyen Trong, Guoxuan Xia et al.CVPR 2025
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan et al.NeurIPS 2025 · 5 citations
- To Smooth or Not? When Label Smoothing Meets Noisy LabelsJiaheng Wei, Hangyu Liu, Tongliang Liu, Gang Niu et al.ICML 2022 · 104 citations
- From Label Smoothing to Label RelaxationJulian Lienen, Eyke HüllermeierAAAI 2021 · 65 citations
- Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion AttacksLukas Struppek, Dominik Hintersdorf, Kristian KerstingICLR 2024 · 26 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
