Language Modeling, Lexical Translation, Reordering: The Training Process of NMT through the Lens of Classical SMT
Elena Voita, Rico Sennrich, Ivan Titov
Abstract
Differently from the traditional statistical MT that decomposes the translation task into distinct separately learned components, neural machine translation uses a single neural network to model the entire translation process. Despite neural machine translation being defacto standard, it is still not clear how NMT models acquire different competences over the course of training, and how this mirrors the different models in traditional SMT. In this work, we look at the competences related to three core SMT components and find that during training, NMT first focuses on learning targetside language modeling, then improves translation quality approaching word-by-word translation, and finally learns more complicated reordering patterns. We show that this behavior holds for several models and language pairs. Additionally, we explain how such an understanding of the training process can be useful in practice and, as an example, show how it can be used to improve vanilla nonautoregressive neural machine translation by guiding teacher model selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano et al.EMNLP 2022 · 18 citations
- Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine TranslationVerna Dankers, Ivan Titov, Dieuwke HupkesEMNLP 2023
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
- Compositional languages emerge in a neural iterated learning modelYi Ren, Shangmin Guo, Matthieu Labeau, Shay B. Cohen et al.ICLR 2020 · 111 citations
- Accurate Word Alignment Induction from Neural Machine TranslationYun Chen, Yang Liu, Guanhua Chen, Xin Jiang et al.EMNLP 2020 · 56 citations
Related papers
- Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering InformationQiu Ran, Yankai Lin, Peng Li, Jie ZhouAAAI 2021 · 82 citations
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 56 citations
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan et al.ACL 2020 · 78 citations
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin et al.AAAI 2020 · 91 citations
- Multi-Granularity Optimization for Non-Autoregressive TranslationYafu Li, Leyang Cui, Yongjing Yin, Yue ZhangEMNLP 2022 · 9 citations
