Distillation Scaling Laws
Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russell Webb
Abstract
We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a teacher needs training. In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size. Conversely, if only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable. Additionally, our large-scale study of distillation increases our understanding of the process and helps inform experimental design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Scaling Laws for Optimal Data MixturesMustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier et al.NeurIPS 2025 · 54 citations
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 25 citations
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
Builds on45
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Towards the Law of Capacity Gap in Distilling Language ModelsChen Zhang, Qiuchi Li, Dawei Song, Zheyu Ye et al.ACL 2025 · 39 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
- Quantifying Cross-Domain Knowledge Distillation in the Presence of Domain ShiftXiangchao Li, Xiao Han, Qing Yang, Xin TongICML 2026
- Supervision Complexity and its Role in Knowledge DistillationHrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim et al.ICLR 2023 · 1 citation
