Lune

NeurIPS2025Top-tier venue

Power Lines: Scaling laws for weight decay and batch size in LLM pre-training

Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness

2025Year
44Citations
17Top-tier citations

Abstract

Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N , dataset size D, and batch size B. Recent work [1] suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N . This law thus provides a method to accurately predict λ opt in advance of large-scale training. We also study scaling laws for optimal batch size B opt (the B enabling lowest loss at a given N, D) and critical batch size B crit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both B opt and B crit scale as power laws in D, independent of model size, N . Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. F Scaling of B opt and B crit : additional details and results F.1 Derivation of "extra data" Eq. ( 6)

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4258bd9a-1314-405f-99a0-c125c8bfdaf9

Cited by top-tier papers17

Ask how each one uses it

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines