Continual Knowledge Distillation for Neural Machine Translation
Yuanchi Zhang, Peng Li, Maosong Sun, Yang Liu
摘要
While many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms. In this work, we propose a method called continual knowledge distillation to take advantage of existing translation models to improve one model of interest. The basic idea is to sequentially transfer knowledge from each trained model to the distilled model. Extensive experiments on Chinese-English and German-English datasets show that our method achieves significant and consistent improvements over strong baselines under both homogeneous and heterogeneous trained model settings and is robust to malicious models. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Finding Sparse Structures for Domain Specific Neural Machine TranslationJianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu 等AAAI 2021 · 被引用 33 次
- Lifelong Language Knowledge DistillationYung-Sung Chuang, Shang-Yu Su, Yun-Nung ChenEMNLP 2020 · 被引用 33 次
- Entropy-Based Vocabulary Substitution for Incremental Learning in Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Yang LiuEMNLP 2022 · 被引用 7 次
- Selective Knowledge Distillation for Neural Machine TranslationFusheng Wang, Jianhao Yan, Fandong Meng, Jie ZhouACL 2021
相关 Paper
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu 等ACL 2024 · 被引用 7 次
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 被引用 40 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng 等AAAI 2020 · 被引用 71 次
