Resolving Discrepancies in Compute-Optimal Scaling of Language Models
Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, Yair Carmon
摘要
Kaplan et al. and Hoffmann et al. developed influential scaling laws for the optimal model size as a function of the compute budget, but these laws yield substantially different predictions. We explain the discrepancy by reproducing the Kaplan scaling law on two datasets (OpenWebText2 and RefinedWeb) and identifying three factors causing the difference: last layer computational cost, warmup duration, and scale-dependent optimizer tuning. With these factors corrected, we obtain excellent agreement with the Hoffmann et al. (i.e.,"Chinchilla") scaling law. Counter to a hypothesis of Hoffmann et al., we find that careful learning rate decay is not essential for the validity of their scaling law. As a secondary result, we derive scaling laws for the optimal learning rate and batch size, finding that tuning the AdamW parameter is essential at lower batch sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal 等NeurIPS 2024 · 被引用 168 次
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 被引用 144 次
- The Art of Scaling Reinforcement Learning Compute for LLMsDevvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal 等ICLR 2026 · 被引用 95 次
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li 等NeurIPS 2025 · 被引用 77 次
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson 等NeurIPS 2025 · 被引用 46 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
相关 Paper
- Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingShane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray 等NeurIPS 2025 · 被引用 44 次
- Predicting Large Model Test Losses with a Noisy Quadratic SystemChuning Li, Chris MaddisonICML 2026
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
- Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMsShane Bergsma, Nolan Simran Dey, Gurpreet Gosal, Gavia Gray 等ICLR 2025 · 被引用 1 次
- Scaling Inference-Efficient Language ModelsSong Bian, Minghao Yan, Shivaram VenkataramanICML 2025
