(Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep Learning
Margaret Li, Sneha Kudugunta, Luke Zettlemoyer
摘要
Modern foundation models rely heavily on using scaling laws to guide crucial training decisions. Researchers often extrapolate the optimal architecture and hyper parameters settings from smaller training runs by describing the relationship between, loss, or task performance, and scale. All components of this process vary, from the specific equation being fit, to the training setup, to the optimization method. Each of these factors may affect the fitted law, and therefore, the conclusions of a given study. We discuss discrepancies in the conclusions that several prior works reach, on questions such as the optimal token to parameter ratio. We augment this discussion with our own analysis of the critical impact that changes in specific details may effect in a scaling study, and the resulting altered conclusions. Additionally, we survey over 50 papers that study scaling trends: while 45 of these papers quantify these trends using a power law, most under-report crucial details needed to reproduce their findings. To mitigate this, we we propose a checklist for authors to consider while contributing to scaling law research.
Researchers have proposed scaling laws to study the scaling of deep learning across multiple domains and for several tasks. Studies of the scaling properties of generalization error with training data size and model capacity predate modern deep learning. Banko & Brill (2001) observed a power law scaling of average validation error on a confusion set disambiguation task with increasing dataset size. The authors also claimed that the model size required to fit a given dataset grows log linearly. As
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Gemstones: A Model Suite for Multi-Faceted Scaling LawsSean McLeish, John Kirchenbauer, David Yu Miller, Siddharth Singh 等NeurIPS 2025 · 被引用 20 次
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and DatasetsMarianna Nezhurina, Tomer Porian, Giovanni Puccetti, Tommie Kerssies 等NeurIPS 2025 · 被引用 8 次
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding 等NeurIPS 2025 · 被引用 4 次
- Scaling Laws of Global Weather ModelsYuejiang Yu, Langwen Huang, Alexandru Calotoiu, Torsten HoeflerICML 2026 · 被引用 4 次
- On the origin of neural scaling laws: from random graphs to natural languageMaissam Barkeshli, Alberto Alfarano, Andrey GromovICML 2026
它引用的顶会 Paper29
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 被引用 315 次
相关 Paper
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
- A Hitchhiker's Guide to Scaling Law EstimationLeshem Choshen, Yang Zhang, Jacob AndreasICML 2025
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
- Broken Neural Scaling LawsEthan Caballero, Kshitij Gupta, Irina Rish, David KruegerICLR 2023 · 被引用 15 次
