Mixed Cross Entropy Loss for Neural Machine Translation
Haoran Li, Wei Lu
Abstract
In neural machine translation, cross entropy (CE) is the standard loss function in two training methods of auto-regressive models, i.e., teacher forcing and scheduled sampling. In this paper, we propose mixed cross entropy loss (mixed CE) as a substitute for CE in both training approaches. In teacher forcing, the model trained with CE regards the translation problem as a one-to-one mapping process, while in mixed CE this process can be relaxed to one-to-many. In scheduled sampling, we show that mixed CE has the potential to encourage the training and testing behaviours to be similar to each other, more effectively mitigating the exposure bias problem. We demonstrate the superiority of mixed CE over CE on several machine translation datasets, WMT'16 Ro-En, WMT'16 Ru-En, and WMT'14 En-De in both teacher forcing and scheduled sampling setups. Furthermore, in WMT'14 En-De, we also find mixed CE consistently outperforms CE on a multi-reference set as well as a challenging paraphrased reference set. We also found the model trained with mixed CE is able to provide a better probability distribution defined over the translation output space. Our code is available at https://github.com/haorannlp/mix .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- GuardT2I: Defending Text-to-Image Models from Adversarial PromptsYijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong et al.NeurIPS 2024 · 74 citations
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 24 citations
- SummAct: Uncovering User Intentions Through Interactive Behaviour SummarisationGuanhua Zhang, Mohamed Adel Naguib Ahmed, Zhiming Hu, Andreas BullingCHI 2025 · 5 citations
- Next-ToBE: Probabilistic Next Token-Bag Exploitation for Activating Anticipatory Capacity in LLMsYihe Liu, Huibin Wang, Xianming Hu, Pinyi Zhang et al.ICLR 2026
- Magmaw: Modality-Agnostic Adversarial Attacks on Machine Learning-Based Wireless Communication SystemsJung-Woo Chang, Ke Sun, Nasimeh Heydaribeni, Seira Hidano et al.NDSS 2025
Builds on1
Related papers
- Scheduled Sampling Based on Decoding Steps for Neural Machine TranslationYijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu et al.EMNLP 2021 · 9 citations
- Understanding and Bridging the Modality Gap for Speech TranslationQingkai Fang, Yang FengACL 2023 · 12 citations
- MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-EntropiesShiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu et al.ACL 2023 · 5 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
