Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation
Yang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang, Chenze Shao
摘要
Although teacher forcing has become the main training paradigm for neural machine translation, it usually makes predictions only conditioned on past information, and hence lacks global planning for the future. To address this problem, we introduce another decoder, called seer decoder, into the encoder-decoder framework during training, which involves future information in target predictions. Meanwhile, we force the conventional decoder to simulate the behaviors of the seer decoder via knowledge distillation. In this way, at test the conventional decoder can perform like the seer decoder without the attendance of it. Experiment results on the Chinese-English, English-German and English-Romanian translation tasks show our method can outperform competitive baselines significantly and achieves greater improvements on the bigger data sets. Besides, the experiments also prove knowledge distillation the best way to transfer knowledge from the seer decoder to the conventional decoder compared to adversarial learning and L2 regularization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reducing Position Bias in Simultaneous Machine Translation with Length-Aware FrameworkShaolei Zhang, Yang FengACL 2022 · 被引用 23 次
- Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?Tianxing He, Jingzhao Zhang, Zhiming Zhou, James R. GlassEMNLP 2021 · 被引用 12 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- EM-Network: Oracle Guided Self-distillation for Sequence LearningJi Won Yoon, Sunghwan Ahn, Hyeonseung Lee, Minchan Kim 等ICML 2023 · 被引用 3 次
它引用的顶会 Paper4
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationChenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng 等AAAI 2020 · 被引用 95 次
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 被引用 44 次
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao 等AAAI 2020 · 被引用 28 次
- Improving Adversarial Text Generation by Modeling the Distant FutureRuiyi Zhang, Changyou Chen, Zhe Gan, Wenlin Wang 等ACL 2020 · 被引用 16 次
相关 Paper
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 被引用 11 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang 等ACL 2022 · 被引用 20 次
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama 等ACL 2020 · 被引用 37 次
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng 等AAAI 2020 · 被引用 71 次
