Lune

NeurIPS2025顶会

Power Lines: Scaling laws for weight decay and batch size in LLM pre-training

Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness

2025年份
44被引次数
17顶会引用

摘要

Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N , dataset size D, and batch size B. Recent work [1] suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N . This law thus provides a method to accurately predict λ opt in advance of large-scale training. We also study scaling laws for optimal batch size B opt (the B enabling lowest loss at a given N, D) and critical batch size B crit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both B opt and B crit scale as power laws in D, independent of model size, N . Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. F Scaling of B opt and B crit : additional details and results F.1 Derivation of "extra data" Eq. ( 6)

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖