Power Lines: Scaling laws for weight decay and batch size in LLM pre-training
Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness
摘要
Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N , dataset size D, and batch size B. Recent work [1] suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N . This law thus provides a method to accurately predict λ opt in advance of large-scale training. We also study scaling laws for optimal batch size B opt (the B enabling lowest loss at a given N, D) and critical batch size B crit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both B opt and B crit scale as power laws in D, independent of model size, N . Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. F Scaling of B opt and B crit : additional details and results F.1 Derivation of "extra data" Eq. ( 6)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf 等ICLR 2026 · 被引用 34 次
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and DurationBruno Mlodozeniec, Pierre Ablin, Louis Béthune, Dan Busbridge 等ICLR 2026 · 被引用 24 次
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model TrainingWilliam Merrill, Shane Arora, Dirk Groeneveld, Hanna HajishirziNeurIPS 2025 · 被引用 23 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- MuLoCo: Muon is a Practical Inner Optimizer for DiLoCoBenjamin Thérien, Xiaolong Huang, Aaron Defazio, Irina Rish 等ICML 2026 · 被引用 15 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
相关 Paper
- How Does Critical Batch Size Scale in Pre-training?Hanlin Zhang, Depen Morwani, Nikhil Vyas, Jingfeng Wu 等ICLR 2025
- Scaling Optimal LR Across Token HorizonsJohan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei 等ICLR 2025
- How to set AdamW's weight decay as you scale model and dataset sizeXi Wang, Laurence AitchisonICML 2025
- Scaling with Collapse: Efficient and Predictable Training of LLM FamiliesShane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad 等ICLR 2026 · 被引用 10 次
- Temporal Scaling Law for Large Language ModelsYizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen 等EMNLP 2025
