Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies
Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, Jingang Wang
Abstract
Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, particularly in large language models, where performance improvements decelerate-a phenomenon known as sub-scaling. This paper revisits these scaling laws by examining the impact of data quality and training strategies on model performance. Through extensive empirical analysis of over 400 models, we identify high data density and non-optimal resource allocation as key factors contributing to subscaling. High data density leads to diminishing returns due to redundant information, while optimal resource allocation is crucial for sustained performance improvements. We propose a sub-optimal scaling law that better predicts performance in sub-scaling regimes, highlighting the importance of data quality and diversity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd5ff377-d976-4781-bc8f-9db400b6352dCited by top-tier papers4
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model PretrainingAnirudh Subramanyam, Yuxin Chen, Robert L. GrossmanICLR 2026 · 6 citations
- Scaling Laws for Code: A More Data-Hungry RegimeXianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang et al.ACL 2026 · 3 citations
- What Scales in Cross-Entropy Scaling Law?Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu et al.ICLR 2026 · 1 citation
- Scaling and Transferability of Annealing Strategies in Large Language Model TrainingSiqi Wang, Zhengyu Chen, Teng Xiao, Zheqi Lv et al.AAAI 2026 · 1 citation
Builds on8
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- Tensor Programs VI: Feature Learning in Infinite Depth Neural NetworksGreg Yang, Dingli Yu, Chen Zhu, Soufiane HayouICLR 2024 · 77 citations
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel et al.ICLR 2024 · 30 citations
- BA-GNN: On Learning Bias-Aware Graph Neural NetworkZhengyu Chen, Teng Xiao, Kun KuangICDE 2022 · 28 citations
Related papers
- How Text Quality Interventions Reshape Neural Scaling Laws for LLMs: Empirical StudyNewsha Ardalani, Feiyang Kang, Michael Kuchnik, Mostafa Elhoushi et al.ICLR 2026
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and RepetitionWeidong Zhou, Fengze Liu, LIU, Ping Guo et al.ICML 2026 · 1 citation
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 144 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan et al.ICLR 2025 · 3 citations
