Selective Knowledge Distillation for Non-Autoregressive Neural Machine Translation
Min Liu, Yu Bao, Chengqi Zhao, Shujian Huang
摘要
Benefiting from the sequence-level knowledge distillation, the Non-Autoregressive Transformer (NAT) achieves great success in neural machine translation tasks. However, existing knowledge distillation has side effects, such as propagating errors from the teacher to NAT students, which may limit further improvements of NAT models and are rarely discussed in existing research. In this paper, we introduce selective knowledge distillation by introducing an NAT evaluator to select NAT-friendly targets that are of high quality and easy to learn. In addition, we introduce a simple yet effective progressive distillation method to boost NAT performance. Experiment results on multiple WMT language directions and several representative NAT models show that our approach can realize a flexible trade-off between the quality and complexity of training data for NAT models, achieving strong performances. Further analysis shows that distilling only 5% of the raw translations can help an NAT outperform its counterpart trained on raw data by about 2.4 BLEU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross 等ICLR 2021 · 被引用 154 次
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin 等AAAI 2020 · 被引用 91 次
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li 等ICML 2022 · 被引用 82 次
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 被引用 56 次
相关 Paper
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ICLR 2021 · 被引用 44 次
- Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationYusheng Liao, Shuyang Jiang, Yiqi Li, Yu Wang 等EMNLP 2023 · 被引用 4 次
- Selective Knowledge Distillation for Neural Machine TranslationFusheng Wang, Jianhao Yan, Fandong Meng, Jie ZhouACL 2021
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ACL 2021
- Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Shuming Shi, Dacheng Tao 等ACL 2022
