Selective Knowledge Distillation for Neural Machine Translation
Fusheng Wang, Jianhao Yan, Fandong Meng, Jie Zhou
摘要
Neural Machine Translation (NMT) models achieve state-of-the-art performance on many translation benchmarks. As an active research field in NMT, knowledge distillation is widely applied to enhance the model's performance by transferring teacher model's knowledge on each training sample. However, previous work rarely discusses the different impacts and connections among these samples, which serve as the medium for transferring teacher knowledge. In this paper, we design a novel protocol that can effectively analyze the different impacts of samples by comparing various samples' partitions. Based on above protocol, we conduct extensive experiments and find that the teacher's knowledge is not the more, the better. Knowledge over specific samples may even hurt the whole performance of knowledge distillation. Finally, to address these issues, we propose two simple yet effective strategies, i.e., batch-level and global-level selections, to pick suitable samples for distillation. We evaluate our approaches on two large-scale machine translation tasks, WMT'14 English-German and WMT'19 Chinese-English. Experimental results show that our approaches yield up to +1.28 and +0.89 BLEU points improvements over the Transformer baseline, respectively. 1 * Equal contribution. † This work was done when Fusheng Wang was interning at Pattern Recognition Center, Wechat AI, Tencent Inc, China. 1 We release our code on https://github.com/Les lieOverfitting/selective distillation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge DistillationThong Thanh Nguyen, Anh Tuan LuuAAAI 2022 · 被引用 46 次
- Matching Structure for Dual LearningHao Fei, Shengqiong Wu, Yafeng Ren, Meishan ZhangICML 2022 · 被引用 44 次
- Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyXiaokai Wei, Sujan Kumar Gonugondla, Shiqi Wang, Wasi Uddin Ahmad 等FSE 2023 · 被引用 29 次
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang 等ACL 2022 · 被引用 20 次
- Shapeshifter: a Parameter-efficient Transformer using Factorized Reshaped MatricesAliakbar Panahi, Seyran Saeedi, Tom ArodzNeurIPS 2021 · 被引用 19 次
它引用的顶会 Paper6
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang 等EMNLP 2020 · 被引用 147 次
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu 等ACL 2020 · 被引用 116 次
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 被引用 97 次
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen 等EMNLP 2020 · 被引用 78 次
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du 等AAAI 2021 · 被引用 44 次
相关 Paper
- Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationMin Liu, Yu Bao, Chengqi Zhao, Shujian HuangAAAI 2023 · 被引用 4 次
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen 等ACL 2023 · 被引用 8 次
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li 等EMNLP 2023 · 被引用 5 次
- Understanding and Improving Lexical Choice in Non-Autoregressive TranslationLiang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong 等ICLR 2021 · 被引用 44 次
- Tailoring Instructions to Student's Learning Levels Boosts Knowledge DistillationYuxin Ren, Zihan Zhong, Xingjian Shi, Yi Zhu 等ACL 2023 · 被引用 7 次
