A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge Distillation
Ayan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy Chakraborty
Abstract
Knowledge distillation (KD) is a technique used to transfer knowledge from a larger "teacher" model into a smaller "student" model. Recent advancements in meta-learning-based knowledge distillation (MetaKD) emphasize that the finetuning of teacher models should be aware of the student's need to achieve better knowledge distillation. However, existing MetaKD methods often lack incentives for the teacher model to improve itself. In this study, we introduce MPDistil, a meta-policy distillation technique, that utilizes novel optimization strategies to foster both collaboration and competition during the fine-tuning of the teacher model in the meta-learning step. Additionally, we propose a curriculum learning framework for the student model in a competitive setup, in which the student model aims to outperform the teacher model by self-training on various tasks. Exhaustive experiments on SuperGLUE and GLUE benchmarks demonstrate the efficacy of MPDistil compared to 20 conventional KD and advanced MetaKD baselines, showing significant performance enhancements in the student model -e.g., a distilled 6-layer BERT model outperforms a 12-layer BERT model on five out of six SuperGLUE tasks. Furthermore, MPDistil, while applied to a large language teacher model (DeBERTa-v2-xxlarge), significantly narrows the performance gap of its smaller student counterpart (DeBERTa-12) by just 4.6% on SuperGLUE. We further demonstrate how higher rewards and customized training curricula strengthen the student model and enhance generalizability.
Published as a conference paper at ICLR 2024 fine-tuning, while the student model is trained to minimise the teacher-student margin. Zhou et al. (2021) pointed out several limitations of this approach. Existing approaches do not optimize the teacher explicitly for the distillation task or consider the student model's capacity during training. This differs from the real-life teacher-student dynamic, in which a teacher continuously enhances her knowledge and teaching skills through repeated student interactions. To address this gap, Zhou et al. (2021) introduced MetaDistil, a meta-learning-based knowledge distillation framework that aims to overcome this limitation by considering the student's predictive performance while fine-tuning the teacher model. In MetaDistil, the teacher model undergoes training on a separate quiz dataset to minimize the knowledge distillation loss between the teacher and student models along with the teacher's task-specific loss. Although MetaDistil addresses several shortcomings of conventional KD methods, it is still far from imitating real-world teacher-student interactions. In a real-world educational context, teachers may not solely focus on narrowing the teacher-student knowledge gap to improve teaching; instead, they aim to maximize the joint performance of students and teachers. Furthermore, in a typical classroom setting, students study various subjects and follow a curriculum to optimize their learning across all subjects. For example, a student might aim to improve her understanding of Physics by studying selected concepts from Mathematics. However, existing knowledge distillation methods tend to distil knowledge for different tasks in isolation, which means they fail to capture the commonalities among tasks. Consequently, the distilled models struggle to generalize effectively to new tasks with limited training data. Another significant challenge with the sole optimization of the teacher-student margin is that it positions the teacher as the "sole role model" for the student without encouraging the student to surpass the teacher's performance.
To address these limitations, we introduce MPDistil, a meta-policy knowledge distillation framework that employs a collaborative learning approach within the context of meta-knowledge distillation 1 . Through empirical analyses, we underscore that optimizing a shared utility function can enhance the teacher's predictive capabilities and ability to impart knowledge to the student model. Furthermore, we establish a meta-reinforcement learning-based "curriculum learning" paradigm in which the student model undergoes fine-tuning on a curriculum -a sequence of tasks. The aim is to enhance the student model's generalization skills, enabling it to surpass the teacher model. This competitive approach empowers the student to outperform the teacher. We assess the effectiveness of MPDistil using two natural language understanding benchmarks, SuperGLUE (Wang et al., 2019) and GLUE (Wang et al., 2018), encompassing 15 NLU tasks. Following the experimental setup adopted by contemporary KD-related studies, we highlight the effectiveness of MPDistil with BERT-base (Devlin et al., 2018) as the teacher model and BERT 6L (with six layers) as the student model. On the SuperGLUE benchmark, the meta update enhances the BERT teacher's performance with a maximum margin of +3%, resulting in a notable improvement in student performance by +5.9
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8e47878-fb67-47dc-8522-c04ae37ababcCited by top-tier papers3
- How to Fine-Tune a Reasoning Model? A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT DataZixian Huang, Kaichen Yang, Xu Huang, Feiyang Hao et al.ICML 2026 · 5 citations
- Heterogeneous Adversarial Play in Interactive EnvironmentsManjie Xu, Xinyi Yang, Jiayu Zhan, Wei Liang et al.NeurIPS 2025 · 4 citations
- Learning from Evolving Training Dynamics: An Entropy-Maximizing Data Curation Strategy for LLM Supervised Post-TrainingMengxiang Zhang, Lingyuan LiuACL 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao et al.AAAI 2023 · 277 citations
Related papers
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 114 citations
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
- f-Divergence Minimization for Sequence-Level Knowledge DistillationYuqiao Wen, Zichao Li, Wenyu Du, Lili MouACL 2023 · 14 citations
