Selective Knowledge Distillation for Neural Machine Translation
Fusheng Wang, Jianhao Yan, Fandong Meng, Jie Zhou
Abstract
Neural Machine Translation (NMT) models achieve state-of-the-art performance on many translation benchmarks. As an active research field in NMT, knowledge distillation is widely applied to enhance the model's performance by transferring teacher model's knowledge on each training sample. However, previous work rarely discusses the different impacts and connections among these samples, which serve as the medium for transferring teacher knowledge. In this paper, we design a novel protocol that can effectively analyze the different impacts of samples by comparing various samples' partitions. Based on above protocol, we conduct extensive experiments and find that the teacher's knowledge is not the more, the better. Knowledge over specific samples may even hurt the whole performance of knowledge distillation. Finally, to address these issues, we propose two simple yet effective strategies, i.e., batch-level and global-level selections, to pick suitable samples for distillation. We evaluate our approaches on two large-scale machine translation tasks, WMT'14 English-German and WMT'19 Chinese-English. Experimental results show that our approaches yield up to +1.28 and +0.89 BLEU points improvements over the Transformer baseline, respectively. 1 * Equal contribution. † This work was done when Fusheng Wang was interning at Pattern Recognition Center, Wechat AI, Tencent Inc, China. 1 We release our code on https://github.com/Les lieOverfitting/selective distillation .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca3dbbf6-ecf2-496c-846a-27490c7b0ba1Cited by top-tier papers17
- Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge DistillationThong Thanh Nguyen, Anh Tuan LuuAAAI 2022 · 46 citations
- Matching Structure for Dual LearningHao Fei, Shengqiong Wu, Yafeng Ren, Meishan ZhangICML 2022 · 44 citations
- Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyXiaokai Wei, Sujan Kumar Gonugondla, Shiqi Wang, Wasi Uddin Ahmad et al.FSE 2023 · 29 citations
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang et al.ACL 2022 · 20 citations
- Shapeshifter: a Parameter-efficient Transformer using Factorized Reshaped MatricesAliakbar Panahi, Seyran Saeedi, Tom ArodzNeurIPS 2021 · 19 citations
Builds on6
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 97 citations
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen et al.EMNLP 2020 · 78 citations
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du et al.AAAI 2021 · 44 citations
Related papers
- Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationMin Liu, Yu Bao, Chengqi Zhao, Shujian HuangAAAI 2023 · 4 citations
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen et al.ACL 2023 · 8 citations
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong et al.ICLR 2021 · 44 citations
- Tailoring Instructions to Student's Learning Levels Boosts Knowledge DistillationYuxin Ren, Zihan Zhong, Xingjian Shi, Yi Zhu et al.ACL 2023 · 7 citations
