Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies
Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, Jingang Wang
摘要
Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, particularly in large language models, where performance improvements decelerate-a phenomenon known as sub-scaling. This paper revisits these scaling laws by examining the impact of data quality and training strategies on model performance. Through extensive empirical analysis of over 400 models, we identify high data density and non-optimal resource allocation as key factors contributing to subscaling. High data density leads to diminishing returns due to redundant information, while optimal resource allocation is crucial for sustained performance improvements. We propose a sub-optimal scaling law that better predicts performance in sub-scaling regimes, highlighting the importance of data quality and diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model PretrainingAnirudh Subramanyam, Yuxin Chen, Robert L. GrossmanICLR 2026 · 被引用 6 次
- Scaling Laws for Code: A More Data-Hungry RegimeXianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang 等ACL 2026 · 被引用 3 次
- What Scales in Cross-Entropy Scaling Law?Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu 等ICLR 2026 · 被引用 1 次
- Scaling and Transferability of Annealing Strategies in Large Language Model TrainingSiqi Wang, Zhengyu Chen, Teng Xiao, Zheqi Lv 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper8
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
- Tensor Programs VI: Feature Learning in Infinite Depth Neural NetworksGreg Yang, Dingli Yu, Chen Zhu, Soufiane HayouICLR 2024 · 被引用 77 次
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel 等ICLR 2024 · 被引用 30 次
- BA-GNN: On Learning Bias-Aware Graph Neural NetworkZhengyu Chen, Teng Xiao, Kun KuangICDE 2022 · 被引用 28 次
相关 Paper
- How Text Quality Interventions Reshape Neural Scaling Laws for LLMs: Empirical StudyNewsha Ardalani, Feiyang Kang, Michael Kuchnik, Mostafa Elhoushi 等ICLR 2026
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and RepetitionWeidong Zhou, Fengze Liu, LIU, Ping Guo 等ICML 2026 · 被引用 1 次
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 被引用 144 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
