From Label Smoothing to Label Relaxation
Julian Lienen, Eyke Hüllermeier
摘要
Regularization of (deep) learning models can be realized at the model, loss, or data level. As a technique somewhere in-between loss and data, label smoothing turns deterministic class labels into probability distributions, for example by uniformly distributing a certain part of the probability mass over all classes. A predictive model is then trained on these distributions as targets, using cross-entropy as loss function. While this method has shown improved performance compared to non-smoothed cross-entropy, we argue that the use of a smoothed though still precise probability distribution as a target can be questioned from a theoretical perspective. As an alternative, we propose a generalized technique called label relaxation, in which the target is a set of probabilities represented in terms of an upper probability distribution. This leads to a genuine relaxation of the target instead of a distortion, thereby reducing the risk of incorporating an undesirable bias in the learning process. Methodically, label relaxation leads to the minimization of a novel type of loss function, for which we propose a suitable closed-form expression for model optimization. The effectiveness of the approach is demonstrated in an empirical study on image data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Credal Self-Supervised LearningJulian Lienen, Eyke HüllermeierNeurIPS 2021 · 被引用 26 次
- Rethinking Image Cropping: Exploring Diverse Compositions from Global ViewsGengyun Jia, Huaibo Huang, Chaoyou Fu, Ran HeCVPR 2022 · 被引用 19 次
- Temporal Label Smoothing for Early Event PredictionHugo Yèche, Alizée Pace, Gunnar Rätsch, Rita KuznetsovaICML 2023 · 被引用 16 次
- Practical Edge Detection via Robust Collaborative LearningYuanbin Fu, Xiaojie GuoACM MM 2023 · 被引用 14 次
- Mitigating Label Noise through Data AmbiguationJulian Lienen, Eyke HüllermeierAAAI 2024 · 被引用 14 次
它引用的顶会 Paper4
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
- Structured Prediction with Partial Labelling through the Infimum LossVivien Cabannes, Alessandro Rudi, Francis R. BachICML 2020 · 被引用 50 次
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
相关 Paper
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 被引用 5 次
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan 等NeurIPS 2025 · 被引用 5 次
- Network as Regularization for Training Deep Neural Networks: Framework, Model and PerformanceKai Tian, Yi Xu, Jihong Guan, Shuigeng ZhouAAAI 2020 · 被引用 6 次
- Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingClara Meister, Elizabeth Salesky, Ryan CotterellACL 2020 · 被引用 4 次
- ACLS: Adaptive and Conditional Label Smoothing for Network CalibrationHyekang Park, Jongyoun Noh, Youngmin Oh, Donghyeon Baek 等ICCV 2023 · 被引用 22 次
