Compact Language Models via Pruning and Knowledge Distillation
Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, Pavlo Molchanov
Abstract
Large language models (LLMs) targeting different deployment scales and sizes are currently produced by training each variant from scratch; this is extremely compute-intensive. In this paper, we investigate if pruning an existing LLM and then re-training it with a fraction (<3%) of the original training data can be a suitable alternative to repeated, full retraining. To this end, we develop a set of practical and effective compression best practices for LLMs that combine depth, width, attention and MLP pruning with knowledge distillation-based retraining; we arrive at these best practices through a detailed empirical exploration of pruning strategies for each axis, methods to combine axes, distillation strategies, and search techniques for arriving at optimal compressed architectures. We use this guide to compress the Nemotron-4 family of LLMs by a factor of 2-4x, and compare their performance to similarly-sized models on a variety of language modeling tasks. Deriving 8B and 4B models from an already pretrained 15B model using our approach requires up to 40x fewer training tokens per model compared to training from scratch; this results in compute cost savings of 1.8x for training the full model family (15B, 8B, and 4B). Minitron models exhibit up to a 16% improvement in MMLU scores compared to training from scratch, perform comparably to other community models such as Mistral 7B, Gemma 7B and Llama-3 8B, and outperform state-of-the-art compression techniques from the literature. We have open-sourced Minitron model weights on Huggingface, with corresponding supplementary material including example code available on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77f66ec9-1fca-429b-aa77-7a368a702958Cited by top-tier papers67
- The Curse of Depth in Large Language ModelsWenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin et al.NeurIPS 2025 · 62 citations
- Towards the Law of Capacity Gap in Distilling Language ModelsChen Zhang, Qiuchi Li, Dawei Song, Zheyu Ye et al.ACL 2025 · 39 citations
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang et al.ICLR 2026 · 24 citations
- Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance AssessmentJun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang et al.AAAI 2025 · 23 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Efficient Hybrid Language Model Compression through Group-Aware SSM PruningAli Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski et al.NeurIPS 2025 · 1 citation
- Pruning Large Language Models with Semi-Structural Adaptive Sparse TrainingWeiyu Huang, Yuezhou Hu, Guohao Jian, Jun Zhu et al.AAAI 2025 · 25 citations
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich et al.NeurIPS 2024 · 72 citations
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 21 citations
- LLaMaFlex: Many-in-one LLMs via Generalized Pruning and Weight SharingRuisi Cai, Saurav Muralidharan, Hongxu Yin, Zhangyang Wang et al.ICLR 2025
