Learning Confidence for Transformer-based Neural Machine Translation
Yu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu, Mu Li
Abstract
Confidence estimation aims to quantify the confidence of the model prediction, providing an expectation of success. A well-calibrated confidence estimate enables accurate failure prediction and proper risk measurement when given noisy samples and out-of-distribution data in real-world settings. However, this task remains a severe challenge for neural machine translation (NMT), where probabilities from softmax distribution fail to describe when the model is probably mistaken. To address this problem, we propose an unsupervised confidence estimate learning jointly with the training of the NMT model. We explain confidence as how many hints the NMT model needs to make a correct prediction, and more hints indicate low confidence. Specifically, the NMT model is given the option to ask for hints to improve translation accuracy at the cost of some slight penalty. Then, we approximate their level of confidence by counting the number of hints the model uses. We demonstrate that our learned confidence estimate achieves high accuracy on extensive sentence/word-level quality estimation tasks. Analytical results verify that our confidence estimate can correctly assess underlying risk in two real-world scenarios: (1) discovering noisy samples and (2) detecting out-of-domain data. We further propose a novel confidence-based instancespecific label smoothing approach based on our learned confidence estimate, which outperforms standard label smoothing 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cecdbd3-e976-40d7-a54c-ea1c00456cf0Cited by top-tier papers6
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
- Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language ModelsAbhishek Kumar, Robert Morabito, Sanzhar Umbet, Jad Kabbara et al.ACL 2024 · 9 citations
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman et al.EMNLP 2024 · 3 citations
- Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language ModelsArtem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Gleb Kuzmin et al.EMNLP 2025 · 1 citation
- Self-Modifying State Modeling for Simultaneous Machine TranslationDonglei Yu, Xiaomian Kang, Yuchen Liu, Yu Zhou et al.ACL 2024
Builds on1
Related papers
- Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding ObjectiveChenze Shao, Fandong Meng, Jiali Zeng, Jie ZhouACL 2024
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu et al.ACL 2023 · 12 citations
- Towards Robust k-Nearest-Neighbor Machine TranslationHui Jiang, Ziyao Lu, Fandong Meng, Chulun Zhou et al.EMNLP 2022 · 16 citations
- Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?Pei Zhang, Baosong Yang, Haoran Wei, Dayiheng Liu et al.EMNLP 2022 · 1 citation
- Prevent the Language Model from being Overconfident in Neural Machine TranslationMengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou et al.ACL 2021
