Power Lines: Scaling laws for weight decay and batch size in LLM pre-training
Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness
Abstract
Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N , dataset size D, and batch size B. Recent work [1] suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N . This law thus provides a method to accurately predict λ opt in advance of large-scale training. We also study scaling laws for optimal batch size B opt (the B enabling lowest loss at a given N, D) and critical batch size B crit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both B opt and B crit scale as power laws in D, independent of model size, N . Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. F Scaling of B opt and B crit : additional details and results F.1 Derivation of "extra data" Eq. ( 6)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4258bd9a-1314-405f-99a0-c125c8bfdaf9Cited by top-tier papers17
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf et al.ICLR 2026 · 34 citations
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and DurationBruno Mlodozeniec, Pierre Ablin, Louis Béthune, Dan Busbridge et al.ICLR 2026 · 24 citations
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model TrainingWilliam Merrill, Shane Arora, Dirk Groeneveld, Hanna HajishirziNeurIPS 2025 · 23 citations
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
- MuLoCo: Muon is a Practical Inner Optimizer for DiLoCoBenjamin Thérien, Xiaolong Huang, Aaron Defazio, Irina Rish et al.ICML 2026 · 15 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
Related papers
- How Does Critical Batch Size Scale in Pre-training?Hanlin Zhang, Depen Morwani, Nikhil Vyas, Jingfeng Wu et al.ICLR 2025
- Scaling Optimal LR Across Token HorizonsJohan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei et al.ICLR 2025
- How to set AdamW's weight decay as you scale model and dataset sizeXi Wang, Laurence AitchisonICML 2025
- Scaling with Collapse: Efficient and Predictable Training of LLM FamiliesShane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad et al.ICLR 2026 · 10 citations
- Temporal Scaling Law for Large Language ModelsYizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen et al.EMNLP 2025
