Non-autoregressive Translation with Layer-Wise Prediction and Deep Supervision
Chenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou, Lei Li
Abstract
How do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient. Recent non-autoregressive translation models speed up the inference, but their quality is still inferior. In this work, we propose DSLP, a highly efficient and high-performance model for machine translation. The key insight is to train a nonautoregressive Transformer with Deep Supervision and feed additional Layer-wise Predictions. We conducted extensive experiments on four translation tasks (both directions of . Results show that our approach consistently improves the BLEU scores compared with respective base models. Specifically, our best variant outperforms the autoregressive model on three translation tasks, while being 14.8 times more efficient in inference. 1 * Work partially done during an internship at ByteDance AI Lab. 1 Our code, training/evaluation scripts, and output are available at https://github.com/chenyangh/DSLP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 219a9f5e-3160-4c08-9872-c5f0feab4356Cited by top-tier papers21
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li et al.ICML 2022 · 82 citations
- On the Learning of Non-Autoregressive TransformersFei Huang, Tianhua Tao, Hao Zhou, Lei Li et al.ICML 2022 · 35 citations
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 26 citations
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 24 citations
- A Character-Level Length-Control Algorithm for Non-Autoregressive Sentence SummarizationPuyuan Liu, Xiang Zhang, Lili MouNeurIPS 2022 · 21 citations
Builds on8
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 181 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Imputer: Sequence Modelling via Imputation and Dynamic ProgrammingWilliam Chan, Chitwan Saharia, Geoffrey E. Hinton, Mohammad Norouzi et al.ICML 2020 · 127 citations
Related papers
- Iterative Refinement in the Continuous Space for Non-Autoregressive Neural Machine TranslationJason Lee, Raphael Shu, Kyunghyun ChoEMNLP 2020 · 20 citations
- RenewNAT: Renewing Potential Translation for Non-autoregressive TransformerPei Guo, Yisheng Xiao, Juntao Li, Min ZhangAAAI 2023 · 9 citations
- Glancing Transformer for Non-Autoregressive Neural Machine TranslationLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang et al.ACL 2021
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch DecoderZhenqiao Song, Hao Zhou, Lihua Qian, Jingjing Xu et al.ICLR 2022 · 13 citations
