On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, Olivier Bachem
Abstract
Knowledge distillation (KD) is widely used for compressing a teacher model to reduce its inference cost and memory footprint, by training a smaller student model. However, current KD methods for auto-regressive sequence models suffer from distribution mismatch between output sequences seen during training and those generated by the student during inference. To address this issue, we introduce Generalized Knowledge Distillation (GKD). Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-generated output sequences by leveraging feedback from the teacher on such sequences. Unlike supervised KD approaches, GKD also offers the flexibility to employ alternative loss functions between the student and teacher, which may be useful when the student lacks the expressivity to mimic the teacher's distribution. Furthermore, GKD facilitates the seamless integration of distillation with RL fine-tuning of language models. We demonstrate the efficacy of GKD for distilling auto-regressive T5 language models for task-specific distillation on summarization, translation, and reasoning tasks, and task-agnostic distillation for instruction tuning. Auto-regressive sequence models, such as language models (LMs), have shown impressive capabilities in numerous tasks, where the key to this success is often scaling the amount of training data as well as * denotes equal contribution. † denotes infrastructure contribution. Correspondence to Rishabh Agarwal
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8950b47-2de1-4f18-b8c3-2c3346d50e72Cited by top-tier papers137
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language ModelsSiyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang et al.ICML 2026 · 245 citations
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski et al.NeurIPS 2024 · 198 citations
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 159 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 138 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
- Online Speculative DecodingXiaoxuan Liu, Lanxiang Hu, Peter Bailis, Alvin Cheung et al.ICML 2024 · 104 citations
Related papers
- DistiLLM: Towards Streamlined Distillation for Large Language ModelsJongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young YunICML 2024 · 86 citations
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu et al.ACL 2024
- f-Divergence Minimization for Sequence-Level Knowledge DistillationYuqiao Wen, Zichao Li, Wenyu Du, Lili MouACL 2023 · 14 citations
- Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved SamplingWenda Xu, Rujun Han, Zifeng Wang, Long T. Le et al.ICLR 2025
- Reinforcement-aware Knowledge Distillation for LLM Reasoningzhaoyang zhang, Shuli Jiang, Yantao Shen, Yuting Zhang et al.ICML 2026
