Improved Natural Language Generation via Loss Truncation
Daniel Kang, Tatsunori Hashimoto
Abstract
Neural language models are usually trained to match the distributional properties of largescale corpora by minimizing the log loss. While straightforward to optimize, this approach forces the model to reproduce all variations in the dataset, including noisy and invalid references (e.g., misannotations and hallucinated facts). Even a small fraction of noisy data can degrade the performance of log loss. As an alternative, prior work has shown that minimizing the distinguishability of generated samples is a principled and robust loss that can handle invalid references. However, distinguishability has not been used in practice due to challenges in optimization and estimation. We propose loss truncation: a simple and scalable procedure which adaptively removes high log loss examples as a way to optimize for distinguishability. Empirically, we demonstrate that loss truncation outperforms existing baselines on distinguishability on a summarization task. Furthermore, we show that samples generated by the loss truncation model have factual accuracy ratings that exceed those of baselines and match human references.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based RetrofittingXinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu et al.AAAI 2024 · 127 citations
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
- Contrastive Learning Reduces Hallucination in ConversationsWeiwei Sun, Zhengliang Shi, Shen Gao, Pengjie Ren et al.AAAI 2023 · 92 citations
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
Builds on3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
Related papers
- Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsTianjian Li, Haoran Xu, Philipp Koehn, Daniel Khashabi et al.ICLR 2024 · 6 citations
- Contrastive Error Attribution for Finetuned Language ModelsFaisal Ladhak, Esin Durmus, Tatsunori HashimotoACL 2023
- Learning with Rejection for Abstractive Text SummarizationMeng Cao, Yue Dong, Jingyi He, Jackie Chi Kit CheungEMNLP 2022 · 10 citations
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
- FACT: Mitigating Inconsistent Hallucinations in LLMs via Fact-Driven Alternating Code-Text TrainingXinxin You, Qixin Sun, Chenwei Yan, Xiao Zhang et al.NeurIPS 2025
