Improving Non-Autoregressive Translation Models Without Distillation
Xiao Shi Huang, Felipe Pérez, Maksims Volkovs
摘要
Transformer-based autoregressive (AR) machine translation models have achieved significant performance improvements, nearing human-level accuracy on some languages. The AR framework translates one token at a time which can be time consuming, especially for long sequences. To accelerate inference, recent work has been exploring non-autoregressive (NAR) approaches that translate blocks of tokens in parallel. Despite significant progress, leading NAR models still lag behind their AR counterparts, and only become competitive when trained with distillation. In this paper we investigate possible reasons behind this performance gap, namely, the indistinguishability of tokens, and mismatch between training and inference. We then propose the Conditional Masked Language Model with Correction (CMLMC) that addresses these problems. Empirically, we show that CMLMC achieves state-of-the-art NAR performance when trained on raw data without distillation and approaches AR performance on multiple datasets. Full code for this work will be released at the time of publication.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper20
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li 等ICML 2022 · 被引用 82 次
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 被引用 26 次
- Accelerating Transformer Inference for Translation via Parallel DecodingAndrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca 等ACL 2023 · 被引用 19 次
- Non-autoregressive Machine Translation with Probabilistic Context-free GrammarShangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang 等NeurIPS 2023 · 被引用 16 次
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li 等AAAI 2023 · 被引用 14 次
相关 Paper
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 被引用 113 次
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross 等ICLR 2021 · 被引用 154 次
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationYusheng Liao, Shuyang Jiang, Yiqi Li, Yu Wang 等EMNLP 2023 · 被引用 4 次
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao 等ACL 2020 · 被引用 58 次
