Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine Translation
Chulun Zhou, Fandong Meng, Jie Zhou, Min Zhang, Hongji Wang, Jinsong Su
摘要
Most dominant neural machine translation (NMT) models are restricted to make predictions only according to the local context of preceding words in a left-to-right manner. Although many previous studies try to incorporate global information into NMT models, there still exist limitations on how to effectively exploit bidirectional global context. In this paper, we propose a Confidence Based Bidirectional Global Context Aware (CB-BGCA) training framework for NMT, where the NMT model is jointly trained with an auxiliary conditional masked language model (CMLM). The training consists of two stages: (1) multi-task joint training; (2) confidence based knowledge distillation. At the first stage, by sharing encoder parameters, the NMT model is additionally supervised by the signal from the CMLM decoder that contains bidirectional global contexts. Moreover, at the second stage, using the CMLM as teacher, we further pertinently incorporate bidirectional global context to the NMT model on its unconfidently-predicted target words via knowledge distillation. Experimental results show that our proposed CB-BGCA training framework significantly improves the NMT model by +1.02, +1.30 and +0.57 BLEU scores on three large-scale translation datasets, namely WMT'14 Englishto-German, WMT'19 Chinese-to-English and WMT'14 English-to-French, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards Better Document-level Relation Extraction via Iterative InferenceLiang Zhang, Jinsong Su, Yidong Chen, Zhongjian Miao 等EMNLP 2022 · 被引用 11 次
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 被引用 10 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain TasksTianyuan Huang, Zepeng Zhu, Hangdi Xing, Zirui Shao 等EMNLP 2025
它引用的顶会 Paper7
- Towards Making the Most of BERT in Neural Machine TranslationJiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao 等AAAI 2020 · 被引用 164 次
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu 等ACL 2020 · 被引用 116 次
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng 等AAAI 2020 · 被引用 71 次
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 被引用 11 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 被引用 55 次
- Guiding Teacher Forcing with Seer Forcing for Neural Machine TranslationYang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang 等ACL 2021
- Multiscale Collaborative Deep Models for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang 等ACL 2020 · 被引用 27 次
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu 等ACL 2022 · 被引用 32 次
- Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine TranslationSongming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen 等ACL 2022 · 被引用 13 次
