From Label Smoothing to Label Relaxation
Julian Lienen, Eyke Hüllermeier
Abstract
Regularization of (deep) learning models can be realized at the model, loss, or data level. As a technique somewhere in-between loss and data, label smoothing turns deterministic class labels into probability distributions, for example by uniformly distributing a certain part of the probability mass over all classes. A predictive model is then trained on these distributions as targets, using cross-entropy as loss function. While this method has shown improved performance compared to non-smoothed cross-entropy, we argue that the use of a smoothed though still precise probability distribution as a target can be questioned from a theoretical perspective. As an alternative, we propose a generalized technique called label relaxation, in which the target is a set of probabilities represented in terms of an upper probability distribution. This leads to a genuine relaxation of the target instead of a distortion, thereby reducing the risk of incorporating an undesirable bias in the learning process. Methodically, label relaxation leads to the minimization of a novel type of loss function, for which we propose a suitable closed-form expression for model optimization. The effectiveness of the approach is demonstrated in an empirical study on image data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 446241d2-6b06-4f5f-92c5-9eee101e0f0fCited by top-tier papers10
- Credal Self-Supervised LearningJulian Lienen, Eyke HüllermeierNeurIPS 2021 · 26 citations
- Rethinking Image Cropping: Exploring Diverse Compositions from Global ViewsGengyun Jia, Huaibo Huang, Chaoyou Fu, Ran HeCVPR 2022 · 19 citations
- Temporal Label Smoothing for Early Event PredictionHugo Yèche, Alizée Pace, Gunnar Rätsch, Rita KuznetsovaICML 2023 · 16 citations
- Practical Edge Detection via Robust Collaborative LearningYuanbin Fu, Xiaojie GuoACM MM 2023 · 14 citations
- Mitigating Label Noise through Data AmbiguationJulian Lienen, Eyke HüllermeierAAAI 2024 · 14 citations
Builds on4
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Structured Prediction with Partial Labelling through the Infimum LossVivien Cabannes, Alessandro Rudi, Francis R. BachICML 2020 · 50 citations
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
Related papers
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan et al.NeurIPS 2025 · 5 citations
- Network as Regularization for Training Deep Neural Networks: Framework, Model and PerformanceKai Tian, Yi Xu, Jihong Guan, Shuigeng ZhouAAAI 2020 · 6 citations
- Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingClara Meister, Elizabeth Salesky, Ryan CotterellACL 2020 · 4 citations
- ACLS: Adaptive and Conditional Label Smoothing for Network CalibrationHyekang Park, Jongyoun Noh, Youngmin Oh, Donghyeon Baek et al.ICCV 2023 · 22 citations
