Weight Distillation: Transferring the Knowledge in Neural Network Parameters
Ye Lin, Yanyang Li, Ziyang Wang, Bei Li, Quan Du, Tong Xiao, Jingbo Zhu
摘要
Knowledge distillation has proven to be effective in model acceleration and compression. It transfers knowledge from a large neural network to a small one by using the large neural network predictions as targets of the small neural network. But this way ignores the knowledge inside the large neural networks, e.g., parameters. Our preliminary study as well as the recent success in pre-training suggests that transferring parameters are more effective in distilling knowledge. In this paper, we propose Weight Distillation to transfer the knowledge in parameters of a large neural network to a small neural network through a parameter generator. On the WMT16 En-Ro, NIST12 Zh-En, and WMT14 En-De machine translation tasks, our experiments show that weight distillation learns a small network that is 1.88∼2.94× faster than the large network but with competitive BLEU performance. When fixing the size of the small networks, weight distillation outperforms knowledge distillation by 0.51∼1.82 BLEU points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Wavelet Knowledge Distillation: Towards Efficient Image-to-Image TranslationLinfeng Zhang, Xin Chen, Xiaobing Tu, Pengfei Wan 等CVPR 2022 · 被引用 105 次
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin 等ICLR 2024 · 被引用 40 次
- Eliciting Knowledge from Large Pre-Trained Models for Unsupervised Knowledge-Grounded ConversationYanyang Li, Jianqiao Zhao, Michael R. Lyu, Liwei WangEMNLP 2022 · 被引用 11 次
- RankNAS: Efficient Neural Architecture Search by Pairwise RankingChi Hu, Chenglong Wang, Xiangnan Ma, Xia Meng 等EMNLP 2021 · 被引用 10 次
- Real-Time Neural Denoising with Render-Aware Knowledge DistillationMengxun Kong, Jie Guo, Chen Wang, Ye Yuan 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper3
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du 等AAAI 2021 · 被引用 44 次
相关 Paper
- Reinforced Multi-Teacher Selection for Knowledge DistillationFei Yuan, Linjun Shou, Jian Pei, Wutao Lin 等AAAI 2021 · 被引用 155 次
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 被引用 3 次
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang 等ACL 2021
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 被引用 4 次
