Lune

NeurIPS2024顶会

Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training

Atli Kosson, Bettina Messmer, Martin Jaggi

2024年份
25被引次数
9顶会引用

摘要

Learning Rate Warmup is a popular heuristic for training neural networks, especially at larger batch sizes, despite limited understanding of its benefits. Warmup decreases the update size Δwt=ηtut\Delta \mathbf{w}_t = \eta_t \mathbf{u}_t early in training by using lower values for the learning rate ηt\eta_t. In this work we argue that warmup benefits training by keeping the overall size of Δwt\Delta \mathbf{w}_t limited, counteracting large initial values of ut\mathbf{u}_t. Focusing on small-scale GPT training with AdamW/Lion, we explore the following question: Why and by which criteria are early updates ut\mathbf{u}_t too large? We analyze different metrics for the update size including the ℓ2\ell_2-norm, resulting directional change, and impact on the representations of the network, providing a new perspective on warmup. In particular, we find that warmup helps counteract large angular updates as well as a limited critical batch size early in training. Finally, we show that the need for warmup can be significantly reduced or eliminated by modifying the optimizer to explicitly normalize ut\mathbf{u}_t based on the aforementioned metrics.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖