Understanding Knowledge Distillation in Non-autoregressive Machine Translation
Chunting Zhou, Jiatao Gu, Graham Neubig
Abstract
Non-autoregressive machine translation (NAT) systems predict a sequence of output tokens in parallel, achieving substantial improvements in generation speed compared to autoregressive models. Existing NAT models usually rely on the technique of knowledge distillation, which creates the training data from a pretrained autoregressive model for better performance. Knowledge distillation is empirically useful, leading to large gains in accuracy for NAT models, but the reason for this success has, as of yet, been unclear. In this paper, we first design systematic experiments to investigate why knowledge distillation is important in NAT training. We find that knowledge distillation can reduce the complexity of data sets and help NAT to model the variations in the output data. Furthermore, a strong correlation is observed between the capacity of an NAT model and the complexity of the distilled data that provides the best translation quality. Based on these findings, we further propose several approaches that can alter the complexity of data sets to improve the performance of NAT models. We achieve state-of-theart performance for NAT-based models, and close the gap with the autoregressive baseline on the WMT14 En-De benchmark. 1 * Equal Contribution. Most work was done during Chunting's internship at FAIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers71
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik et al.NeurIPS 2023 · 212 citations
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 155 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
Builds on2
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
Related papers
- Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationMin Liu, Yu Bao, Chengqi Zhao, Shujian HuangAAAI 2023 · 4 citations
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong et al.ICLR 2021 · 44 citations
- Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationYusheng Liao, Shuyang Jiang, Yiqi Li, Yu Wang et al.EMNLP 2023 · 4 citations
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong et al.ACL 2021
