Non-autoregressive Translation with Layer-Wise Prediction and Deep Supervision
Chenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou, Lei Li
摘要
How do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient. Recent non-autoregressive translation models speed up the inference, but their quality is still inferior. In this work, we propose DSLP, a highly efficient and high-performance model for machine translation. The key insight is to train a nonautoregressive Transformer with Deep Supervision and feed additional Layer-wise Predictions. We conducted extensive experiments on four translation tasks (both directions of . Results show that our approach consistently improves the BLEU scores compared with respective base models. Specifically, our best variant outperforms the autoregressive model on three translation tasks, while being 14.8 times more efficient in inference. 1 * Work partially done during an internship at ByteDance AI Lab. 1 Our code, training/evaluation scripts, and output are available at https://github.com/chenyangh/DSLP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li 等ICML 2022 · 被引用 82 次
- On the Learning of Non-Autoregressive TransformersFei Huang, Tianhua Tao, Hao Zhou, Lei Li 等ICML 2022 · 被引用 35 次
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 被引用 26 次
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 被引用 24 次
- A Character-Level Length-Control Algorithm for Non-Autoregressive Sentence SummarizationPuyuan Liu, Xiang Zhang, Lili MouNeurIPS 2022 · 被引用 21 次
它引用的顶会 Paper8
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 被引用 181 次
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross 等ICLR 2021 · 被引用 154 次
- Imputer: Sequence Modelling via Imputation and Dynamic ProgrammingWilliam Chan, Chitwan Saharia, Geoffrey E. Hinton, Mohammad Norouzi 等ICML 2020 · 被引用 127 次
相关 Paper
- Iterative Refinement in the Continuous Space for Non-Autoregressive Neural Machine TranslationJason Lee, Raphael Shu, Kyunghyun ChoEMNLP 2020 · 被引用 20 次
- RenewNAT: Renewing Potential Translation for Non-autoregressive TransformerPei Guo, Yisheng Xiao, Juntao Li, Min ZhangAAAI 2023 · 被引用 9 次
- Glancing Transformer for Non-Autoregressive Neural Machine TranslationLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang 等ACL 2021
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 被引用 125 次
- switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch DecoderZhenqiao Song, Hao Zhou, Lihua Qian, Jingjing Xu 等ICLR 2022 · 被引用 13 次
