Revisiting Neural Scaling Laws in Language and Vision
Ibrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua Zhai
摘要
The remarkable progress in deep learning in recent years is largely driven by improvements in scale, where bigger models are trained on larger datasets for longer schedules. To predict the benefit of scale empirically, we argue for a more rigorous methodology based on the extrapolation loss, instead of reporting the bestfitting (interpolating) parameters. We then present a recipe for estimating scaling law parameters reliably from learning curves. We demonstrate that it extrapolates more accurately than previous methods in a wide range of architecture families across several domains, including image classification, neural machine translation (NMT) and language modeling, in addition to tasks from the BIG-Bench evaluation benchmark. Finally, we release a benchmark dataset comprising of 90 evaluation tasks to facilitate research in this domain. Related work Power law scaling in deep neural architectures has been verified in a wide range of domains, including image classification [2, 20, 32, 40] , language modeling [20, 23, 32] , NMT [3, [18] [19] [20] , and speech recognition [20] . To explain this theoretically, at least for data scaling, several works have argued
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper79
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu 等NeurIPS 2024 · 被引用 125 次
- Getting ViT in Shape: Scaling Laws for Compute-Optimal Model DesignIbrahim M. Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, Lucas BeyerNeurIPS 2023 · 被引用 122 次
- Symbolic Regression with a Learned Concept LibraryArya Grayeli, Atharva Sehgal, Omar Costilla-Reyes, Miles D. Cranmer 等NeurIPS 2024 · 被引用 105 次
- Generalization on the Unseen, Logic Reasoning and Degree CurriculumEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Kevin RizkICML 2023 · 被引用 68 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
相关 Paper
- A Hitchhiker's Guide to Scaling Law EstimationLeshem Choshen, Yang Zhang, Jacob AndreasICML 2025
- Broken Neural Scaling LawsEthan Caballero, Kshitij Gupta, Irina Rish, David KruegerICLR 2023 · 被引用 15 次
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman 等ICLR 2026 · 被引用 8 次
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding 等NeurIPS 2025 · 被引用 4 次
