Prevent the Language Model from being Overconfident in Neural Machine Translation
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou
摘要
The Neural Machine Translation (NMT) model is essentially a joint language model conditioned on both the source sentence and partial translation. Therefore, the NMT model naturally involves the mechanism of the Language Model (LM) that predicts the next token only based on partial translation. Despite its success, NMT still suffers from the hallucination problem, generating fluent but inadequate translations. The main reason is that NMT pays excessive attention to the partial translation while neglecting the source sentence to some extent, namely overconfidence of the LM. Accordingly, we define the Margin between the NMT and the LM, calculated by subtracting the predicted probability of the LM from that of the NMT model for each token. The Margin is negatively correlated to the overconfidence degree of the LM. Based on the property, we propose a Margin-based Token-level Objective (MTO) and a Margin-based Sentencelevel Objective (MSO) to maximize the Margin for preventing the LM from being overconfident. Experiments on WMT14 Englishto-German, WMT19 Chinese-to-English, and WMT14 English-to-French translation tasks demonstrate the effectiveness of our approach, with 1.36, 1.50, and 0.63 BLEU improvements, respectively, compared to the Transformer baseline. The human evaluation further verifies that our approaches improve translation adequacy as well as fluency. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesRuslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen 等NeurIPS 2024 · 被引用 70 次
- Towards Improving Faithfulness in Abstractive SummarizationXiuying Chen, Mingzhe Li, Xin Gao, Xiangliang ZhangNeurIPS 2022 · 被引用 39 次
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang 等ACL 2022 · 被引用 20 次
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang 等AAAI 2023 · 被引用 19 次
它引用的顶会 Paper5
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao 等AAAI 2020 · 被引用 28 次
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 被引用 21 次
- Towards Enhancing Faithfulness for Neural Machine TranslationRongxiang Weng, Heng Yu, Xiangpeng Wei, Weihua LuoEMNLP 2020 · 被引用 18 次
- Parallel Corpus Filtering via Pre-trained Language ModelsBoliang Zhang, Ajay Nagesh, Kevin KnightACL 2020 · 被引用 17 次
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 被引用 11 次
相关 Paper
- Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding ObjectiveChenze Shao, Fandong Meng, Jiali Zeng, Jie ZhouACL 2024
- Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterDavid Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussàACL 2023 · 被引用 25 次
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng 等EMNLP 2020 · 被引用 32 次
- M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine TranslationHao Wang, Linlong Xu, Heng Liu, Yangyang Liu 等ACL 2026
- Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang 等ACL 2022 · 被引用 32 次
