Adaptive Label Smoothing with Self-Knowledge in Natural Language Generation
Dongkyu Lee, Ka Chun Cheung, Nevin L. Zhang
Abstract
Overconfidence has been shown to impair generalization and calibration of a neural network. Previous studies remedy this issue by adding a regularization term to a loss function, preventing a model from making a peaked distribution. Label smoothing smoothes target labels with a pre-defined prior label distribution; as a result, a model is learned to maximize the likelihood of predicting the soft label. Nonetheless, the amount of smoothing is the same in all samples and remains fixed in training. In other words, label smoothing does not reflect the change in probability distribution mapped by a model over the course of training. To address this issue, we propose a regularization scheme that brings dynamic nature into the smoothing parameter by taking model probability distribution into account, thereby varying the parameter per instance. A model in training self-regulates the extent of smoothing on the fly during forward propagation. Furthermore, inspired by recent work in bridging label smoothing and knowledge distillation, our work utilizes self-knowledge as a prior label distribution in softening target labels, and presents theoretical support for the regularization effect by knowledge distillation and the dynamic smoothing parameter. Our regularizer is validated comprehensively, and the result illustrates marked improvements in model generalization and calibration, enhancing robustness and trustworthiness of a model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2428363-df2e-4a28-b505-9dc71ba35aa0Cited by top-tier papers4
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan et al.NeurIPS 2025 · 5 citations
- Training High Performance Spiking Neural Network by Temporal Model CalibrationJiaqi Yan, Changping Wang, De Ma, Huajin Tang et al.ICML 2025
- Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingTianci Liu, Ruirui Li, Zihan Dong, Hui Liu et al.ICML 2025
- MAFA: Managing False Negatives for Vision-Language Pre-TrainingJaeseok Byun, Dohoon Kim, Taesup MoonCVPR 2024
Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Self-Knowledge Distillation with Progressive Refinement of TargetsKyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum HwangICCV 2021 · 251 citations
- Learning Better Structured Representations Using Low-rank Adaptive Label SmoothingAsish Ghoshal, Xilun Chen, Sonal Gupta, Luke Zettlemoyer et al.ICLR 2021 · 16 citations
- Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingClara Meister, Elizabeth Salesky, Ryan CotterellACL 2020 · 4 citations
Related papers
- Training Deep Neural Networks with Virtual Smoothing ClassesZhiyang Zhou, Siwei Wei, Xudong Zhang, Wensheng Dou et al.AAAI 2025
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang et al.CVPR 2020
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
- From Label Smoothing to Label RelaxationJulian Lienen, Eyke HüllermeierAAAI 2021 · 65 citations
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou et al.ICLR 2021 · 209 citations
