Self-Improvement of Non-autoregressive Model via Sequence-Level Distillation
Yusheng Liao, Shuyang Jiang, Yiqi Li, Yu Wang, Yanfeng Wang
Abstract
Although Non-autoregressive Transformer (NAT) models have achieved great success in terms of fast inference speed, this speedup comes with a performance drop due to the inherent multi-modality problem of the NAT model. Previous works commonly alleviate this problem by replacing the target side of the raw data with distilled data generated by Autoregressive Transformer (AT) models. However, the multi-modality problem in the distilled data is still significant and thus limits further improvement of the NAT models. In this paper, we propose a method called Sequence-Level Self-Distillation (SLSD), which aims to generate distilled data by the NAT model itself, eliminating the need for additional teacher networks. Furthermore, SLSD can adapt to different NAT models without precise adjustments since the self-distilled data is generated from the same types of NAT models. We conduct extensive experiments on WMT14 EN↔DE and WMT16 EN↔RO and choose four classic NAT models as the backbones to validate the generality and effectiveness of SLSD. The results show that our approach can consistently improve all models on both raw data and distilled data without sacrificing the inference speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0997b22-c12a-49a8-ba64-0b29baad2b29Builds on11
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
- Improving Non-Autoregressive Translation Models Without DistillationXiao Shi Huang, Felipe Pérez, Maksims VolkovsICLR 2022 · 60 citations
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
Related papers
- Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationMin Liu, Yu Bao, Chengqi Zhao, Shujian HuangAAAI 2023 · 4 citations
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong et al.ICLR 2021 · 44 citations
- Rephrasing the Reference for Non-autoregressive Machine TranslationChenze Shao, Jinchao Zhang, Jie Zhou, Yang FengAAAI 2023 · 6 citations
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li et al.ICML 2022 · 82 citations
