Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing
Asish Ghoshal, Xilun Chen, Sonal Gupta, Luke Zettlemoyer, Yashar Mehdad
Abstract
Training with soft targets instead of hard targets has been shown to improve performance and calibration of deep neural networks. Label smoothing is a popular way of computing soft targets, where one-hot encoding of a class is smoothed with a uniform distribution. Owing to its simplicity, label smoothing has found wide-spread use for training deep neural networks on a wide variety of tasks, ranging from image and text classification to machine translation and semantic parsing. Complementing recent empirical justification for label smoothing, we obtain PAC-Bayesian generalization bounds for label smoothing and show that the generalization error depends on the choice of the noise (smoothing) distribution. Then we propose low-rank adaptive label smoothing (LORAS): a simple yet novel method for training with learned soft targets that generalizes label smoothing and adapts to the latent structure of the label space in structured prediction tasks. Specifically, we evaluate our method on semantic parsing tasks and show that training with appropriately smoothed soft targets can significantly improve accuracy and model calibration, especially in low-resource settings. Used in conjunction with pre-trained sequence-to-sequence models, our method achieves state of the art performance on four semantic parsing data sets. LORAS can be used with any model, improves performance and implicit model calibration without increasing the number of model parameters, and can be scaled to problems with large label spaces containing tens of thousands of labels.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- A Gift from Label Smoothing: Robust Training with Adaptive Label Smoothing via Auxiliary Classifier under Label NoiseJongwoo Ko, Bongsoo Yi, Se-Young YunAAAI 2023 · 9 citations
- Language Modelling via Learning to RankArvid Frydenlund, Gagandeep Singh, Frank RudziczAAAI 2022 · 9 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- Aligned Objective for Soft-Pseudo-Label Generation in Supervised LearningNing Xu, Yihao Hu, Congyu Qiao, Xin GengICML 2024 · 1 citation
- Hard Gate Knowledge Distillation - Leverage Calibration for Robust and Reliable Language ModelDongkyu Lee, Zhiliang Tian, Yingxiu Zhao, Ka Chun Cheung et al.EMNLP 2022
Related papers
- Class Adaptive Network CalibrationBingyuan Liu, Jérôme Rony, Adrian Galdran, Jose Dolz et al.CVPR 2023
- From Label Smoothing to Label RelaxationJulian Lienen, Eyke HüllermeierAAAI 2021 · 65 citations
- Neutral residues: revisiting adapters for model extensionFranck Signe Talla, Edouard Grave, Hervé JégouICML 2025
- Adaptive Utilization of Low-Rank Adaptation via Conditioned GatingGuang Yang, Changhao Guan, Chao Huang, Yufeng Chen et al.ICML 2026
- Flat-LoRA: Low-Rank Adaptation over a Flat Loss LandscapeTao Li, Zhengbao He, Yujun Li, Yasheng Wang et al.ICML 2025
