Lune

NeurIPS2024Top-tier venue

Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training

Atli Kosson, Bettina Messmer, Martin Jaggi

2024Year
25Citations
9Top-tier citations

Abstract

Learning Rate Warmup is a popular heuristic for training neural networks, especially at larger batch sizes, despite limited understanding of its benefits. Warmup decreases the update size Δwt=ηtut\Delta \mathbf{w}_t = \eta_t \mathbf{u}_t early in training by using lower values for the learning rate ηt\eta_t. In this work we argue that warmup benefits training by keeping the overall size of Δwt\Delta \mathbf{w}_t limited, counteracting large initial values of ut\mathbf{u}_t. Focusing on small-scale GPT training with AdamW/Lion, we explore the following question: Why and by which criteria are early updates ut\mathbf{u}_t too large? We analyze different metrics for the update size including the ℓ2\ell_2-norm, resulting directional change, and impact on the representations of the network, providing a new perspective on warmup. In particular, we find that warmup helps counteract large angular updates as well as a limited critical batch size early in training. Finally, we show that the need for warmup can be significantly reduced or eliminated by modifying the optimizer to explicitly normalize ut\mathbf{u}_t based on the aforementioned metrics.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 81e1b115-0ace-41a3-92d3-cae9eca5eba9

Cited by top-tier papers9

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines