MaxSup: Overcoming Representation Collapse in Label Smoothing
Yuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan, Yifei Dong, Mario Fritz, Margret Keuper
Abstract
Label Smoothing (LS) is widely adopted to reduce overconfidence in neural network predictions and improve generalization. Despite these benefits, recent studies reveal two critical issues with LS. First, LS induces overconfidence in misclassified samples. Second, it compacts feature representations into overly tight clusters, diluting intra-class diversity, although the precise cause of this phenomenon remained elusive. In this paper, we analytically decompose the LS-induced loss, exposing two key terms: (i) a regularization term that dampens overconfidence only when the prediction is correct, and (ii) an error-amplification term that arises under misclassifications. This latter term compels the network to reinforce incorrect predictions with undue certainty, exacerbating representation collapse. To address these shortcomings, we propose Max Suppression (MaxSup), which applies uniform regularization to both correct and incorrect predictions by penalizing the top-1 logit rather than the ground-truth logit. Through extensive feature-space analyses, we show that MaxSup restores intra-class variation and sharpens inter-class boundaries. Experiments on large-scale image classification and multiple downstream tasks confirm that MaxSup is a more robust alternative to LS. Code is available at: https://github.com/ZhouYuxuanYX/Maximum-Suppression-Regularization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2e983ab-4f4e-4324-a86e-80049aa36b9bCited by top-tier papers2
- Flatter Tokens are More Valuable for Speculative Draft Model TrainingJiaming Fan, Daming Cao, Xiangzhong Luo, Jiale Fu et al.ICLR 2026 · 2 citations
- Your Dissimilarities Define You: Complementary Learning Exploiting Class DiversitiesDimitrios Katsikas, Nikolaos Passalis, Anastasios TefasCVPR 2026
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
- HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningShiming Chen, Guo-Sen Xie, Yang Liu, Qinmu Peng et al.NeurIPS 2021 · 190 citations
Related papers
- Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix ItGuoxuan Xia, Olivier Laurent, Gianni Franchi, Christos-Savvas BouganisICLR 2025
- Training Deep Neural Networks with Virtual Smoothing ClassesZhiyang Zhou, Siwei Wei, Xudong Zhang, Wensheng Dou et al.AAAI 2025
- From Label Smoothing to Label RelaxationJulian Lienen, Eyke HüllermeierAAAI 2021 · 65 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- Adversarial Unlearning: Reducing Confidence Along Adversarial DirectionsAmrith Setlur, Benjamin Eysenbach, Virginia Smith, Sergey LevineNeurIPS 2022 · 26 citations
