Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine Translation
Chulun Zhou, Fandong Meng, Jie Zhou, Min Zhang, Hongji Wang, Jinsong Su
Abstract
Most dominant neural machine translation (NMT) models are restricted to make predictions only according to the local context of preceding words in a left-to-right manner. Although many previous studies try to incorporate global information into NMT models, there still exist limitations on how to effectively exploit bidirectional global context. In this paper, we propose a Confidence Based Bidirectional Global Context Aware (CB-BGCA) training framework for NMT, where the NMT model is jointly trained with an auxiliary conditional masked language model (CMLM). The training consists of two stages: (1) multi-task joint training; (2) confidence based knowledge distillation. At the first stage, by sharing encoder parameters, the NMT model is additionally supervised by the signal from the CMLM decoder that contains bidirectional global contexts. Moreover, at the second stage, using the CMLM as teacher, we further pertinently incorporate bidirectional global context to the NMT model on its unconfidently-predicted target words via knowledge distillation. Experimental results show that our proposed CB-BGCA training framework significantly improves the NMT model by +1.02, +1.30 and +0.57 BLEU scores on three large-scale translation datasets, namely WMT'14 Englishto-German, WMT'19 Chinese-to-English and WMT'14 English-to-French, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Towards Better Document-level Relation Extraction via Iterative InferenceLiang Zhang, Jinsong Su, Yidong Chen, Zhongjian Miao et al.EMNLP 2022 · 11 citations
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 10 citations
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen et al.ACL 2023 · 8 citations
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 3 citations
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain TasksTianyuan Huang, Zepeng Zhu, Hangdi Xing, Zirui Shao et al.EMNLP 2025
Builds on7
- Towards Making the Most of BERT in Neural Machine TranslationJiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao et al.AAAI 2020 · 164 citations
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng et al.AAAI 2020 · 71 citations
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 11 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 55 citations
- Guiding Teacher Forcing with Seer Forcing for Neural Machine TranslationYang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang et al.ACL 2021
- Multiscale Collaborative Deep Models for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang et al.ACL 2020 · 27 citations
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine TranslationSongming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen et al.ACL 2022 · 13 citations
