Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine Translation
Junliang Guo, Linli Xu, Enhong Chen
Abstract
The masked language model has received remarkable attention due to its effectiveness on various natural language processing tasks. However, few works have adopted this technique in the sequence-to-sequence models. In this work, we introduce a jointly masked sequence-to-sequence model and explore its application on non-autoregressive neural machine translation (NAT). Specifically, we first empirically study the functionalities of the encoder and the decoder in NAT models, and find that the encoder takes a more important role than the decoder regarding the translation quality. Therefore, we propose to train the encoder more rigorously by masking the encoder input while training. As for the decoder, we propose to train it based on the consecutive masking of the decoder input with an n-gram loss function to alleviate the problem of translating duplicate words. The two types of masks are applied to the model jointly at the training stage. We conduct experiments on five benchmark machine translation tasks, and our model can achieve 27.69/32.24 BLEU scores on WMT14 English-German/German-English tasks with 5+ times speed up compared with an autoregressive model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0709e150-1187-479a-8411-343eb2aa73b5Cited by top-tier papers27
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik et al.NeurIPS 2023 · 212 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Step-unrolled Denoising Autoencoders for Text GenerationNikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen et al.ICLR 2022 · 142 citations
- Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering InformationQiu Ran, Yankai Lin, Peng Li, Jie ZhouAAAI 2021 · 82 citations
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li et al.ICML 2022 · 82 citations
Builds on3
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationChenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng et al.AAAI 2020 · 95 citations
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin et al.AAAI 2020 · 91 citations
- IntroVNMT: An Introspective Model for Variational Neural Machine TranslationXin Sheng, Linli Xu, Junliang Guo, Jingchang Liu et al.AAAI 2020 · 6 citations
Related papers
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationYong Wang, Shilin He, Guanhua Chen, Yun Chen et al.EMNLP 2022 · 4 citations
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei et al.NeurIPS 2020 · 72 citations
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li et al.AAAI 2023 · 14 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
