The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
Tian Jin, Ahmed Imtiaz Humayun, Utku Evci, Suvinay Subramanian, Amir Yazdanbakhsh, Dan Alistarh, Gintare Karolina Dziugaite
Abstract
Pruning eliminates unnecessary parameters in neural networks; it offers a promising solution to the growing computational demands of large language models (LLMs). While many focus on post-training pruning, sparse pre-training-which combines pruning and pre-training into a single phase-provides a simpler alternative. In this work, we present the first systematic exploration of optimal sparse pretraining configurations for LLMs through an examination of 80 unique pruning schedules across different sparsity levels and training durations. We find that initiating pruning at 25% of total training compute and concluding at 75% achieves near-optimal final evaluation loss. These findings provide valuable insights for efficient and effective sparse pre-training of LLMs. Furthermore, we propose a new scaling law that modifies the Chinchilla scaling law to use the average parameter count over pre-training. Through empirical and theoretical validation, we demonstrate that this modified scaling law accurately models evaluation loss for both sparsely and densely pre-trained LLMs, unifying scaling laws across pretraining paradigms. Our findings indicate that while sparse pre-training achieves the same final model quality as dense pre-training for equivalent compute budgets, it provides substantial benefits through reduced model size, enabling significant potential computational savings during inference. INTRODUCTION As language models grow in size and train on more data, they consistently demonstrate improved performance (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- QuEST: Stable Training of LLMs with 1-Bit Weights and ActivationsAndrei Panferov, Jiale Chen, Soroush Tabesh, Mahdi Nikdan et al.ICML 2025
- When Data Is Scarce: Scaling Sparse Language Models with Repeated TrainingBoqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal et al.ICML 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 144 citations
- EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language ModelsXingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao et al.ACL 2026
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model PretrainingAnirudh Subramanyam, Yuxin Chen, Robert L. GrossmanICLR 2026 · 6 citations
- P² Law: Scaling Law for Post-Training After Model PruningXiaodong Chen, Yuxuan Hu, Xiaokang Zhang, Yanling Wang et al.ACL 2025
- Predicting Large Model Test Losses with a Noisy Quadratic SystemChuning Li, Chris MaddisonICML 2026
