Quantification of Large Language Model Distillation
Sunbowen Lee, Junting Zhou, Chang Ao, Kaige Li, Xeron Du, Sirui He, Haihong Wu, Tianci Liu, Jiaheng Liu, Hamid Alinejad-Rokny, Min Yang, Yitao Liang
Abstract
Model distillation is a fundamental technique in building large language models (LLMs), transferring knowledge from a teacher model to a student model. However, distillation can lead to model homogenization, reducing diversity among models and impairing their ability to robustly handle complex or novel tasks. These limitations underscore the need to systematically quantify the distillation process and its impact. In this work, we propose a framework to evaluate and quantify model distillation. Our method addresses two key aspects: (1) Identifying identity cognition contradictions to assess discrepancies in how models perceive and represent identity-related information, and (2) Analyzing multi-granularity response similarities across models to measure the extent of homogenization. Experimental results demonstrate two key insights: (1) Well-known closedsource and open-source LLMs usually exhibit high distillation degrees, except for Claude, Doubao, and Gemini. (2) Base LLMs show higher distillation degrees compared to aligned LLMs. By offering a systematic approach to improve the transparency of LLM data distillation, we call for LLMs with more independent development and more transparent technical reports to improve LLMs' robustness and safety. The code and data are available at https://github.com/Aegis1863/LLMs-Distillation-Quantification .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext daa9a481-8273-40fc-be59-7311873cd4a6Cited by top-tier papers5
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo et al.ICLR 2026 · 21 citations
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy ModelsJunhao Liu, Haonan Yu, Zhenyu Yan, Xin ZhangACL 2026 · 2 citations
- When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use BehaviorsChenghao Yang, Yuning Zhang, Zhoufutu Wen, Tao Gong et al.ACL 2026 · 1 citation
- Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningZhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao et al.ACL 2026
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
Related papers
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen et al.EMNLP 2024 · 1 citation
- Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation PipelinesKatherine Lambert, Sasha LuccioniICML 2026 · 1 citation
- Idiosyncrasies in Large Language ModelsMingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter et al.ICML 2025
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar et al.ICLR 2024 · 39 citations
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
