BERT Learns to Teach: Knowledge Distillation with Meta Learning
Wangchunshu Zhou, Canwen Xu, Julian J. McAuley
Abstract
We present Knowledge Distillation with Meta Learning (MetaDistil), a simple yet effective alternative to traditional knowledge distillation (KD) methods where the teacher model is fixed during training. We show the teacher network can learn to better transfer knowledge to the student network (i.e., learning to teach) with the feedback from the performance of the distilled student network in a meta learning framework. Moreover, we introduce a pilot update mechanism to improve the alignment between the inner-learner and meta-learner in meta learning algorithms that focus on an improved inner-learner. Experiments on various benchmarks show that MetaDistil can yield significant improvements compared with traditional KD algorithms and is less sensitive to the choice of different student capacity and hyperparameters, facilitating the use of KD on different tasks and models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49b5d146-4092-452f-9def-b007d6a06744Cited by top-tier papers23
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
- Towards the Law of Capacity Gap in Distilling Language ModelsChen Zhang, Qiuchi Li, Dawei Song, Zheyu Ye et al.ACL 2025 · 39 citations
- PROD: Progressive Distillation for Dense RetrievalZhenghao Lin, Yeyun Gong, Xiao Liu, Hang Zhang et al.WWW 2023 · 33 citations
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie et al.AAAI 2025 · 31 citations
Builds on11
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
Related papers
- A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge DistillationAyan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy ChakrabortyICLR 2024 · 16 citations
- Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingMingi Ji, Byeongho Heo, Sungrae ParkAAAI 2021 · 194 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge DistillationGaurav Patel, Konda Reddy Mopuri, Qiang QiuCVPR 2023
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
