MixKD: Towards Efficient Distillation of Large-scale Language Models
Kevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou, Weizhu Chen, Changyou Chen, Lawrence Carin
Abstract
Large-scale language models have recently demonstrated impressive empirical performance. Nevertheless, the improved results are attained at the price of bigger models, more power consumption, and slower inference, which hinder their applicability to low-resource (both memory and computation) platforms. Knowledge distillation (KD) has been demonstrated as an effective framework for compressing such big models. However, large-scale neural network systems are prone to memorize training instances, and thus tend to make inconsistent predictions when the data distribution is altered slightly. Moreover, the student model has few opportunities to request useful information from the teacher model when there is limited task-specific data available. To address these issues, we propose MixKD, a data-agnostic distillation framework that leverages mixup, a simple yet efficient data augmentation approach, to endow the resulting model with stronger generalization ability. Concretely, in addition to the original training examples, the student model is encouraged to mimic the teacher's behavior on the linear interpolation of example pairs as well. We prove from a theoretical perspective that under reasonable conditions MixKD gives rise to a smaller gap between the generalization error and the empirical error. To verify its effectiveness, we conduct experiments on the GLUE benchmark, where MixKD consistently leads to significant gains over the standard KD training, and outperforms several competitive baselines. Experiments under a limited-data setting and ablation studies further demonstrate the advantages of the proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 64 citations
- Dynamic Knowledge Distillation for Pre-trained Language ModelsLei Li, Yankai Lin, Shuhuai Ren, Peng Li et al.EMNLP 2021 · 32 citations
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationGaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. LaradjiEMNLP 2023 · 11 citations
Builds on6
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- GraphMix: Improved Training of GNNs for Semi-Supervised LearningVikas Verma, Meng Qu, Kenji Kawaguchi, Alex Lamb et al.AAAI 2021 · 157 citations
- Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox ModelDongdong Wang, Yandong Li, Liqiang Wang, Boqing GongCVPR 2020
Related papers
- Augmentation with Projection: Towards an Effective and Efficient Data Augmentation Paradigm for DistillationZiqi Wang, Yuexin Wu, Frederick Liu, Daogao Liu et al.ICLR 2023 · 2 citations
- Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingAhmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi RezagholizadehEMNLP 2021 · 25 citations
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge DistillationZihao Tang, Zheqi Lv, Shengyu Zhang, Yifan Zhou et al.ICLR 2024 · 5 citations
- AMiD: Knowledge Distillation for LLMs with α-mixture Assistant DistributionDonghyeok Shin, Yeongmin Kim, Suhyeon Jo, Byeonghu Na et al.ICLR 2026 · 3 citations
