Lune

NeurIPS2024顶会

Why Warmup the Learning Rate? Underlying Mechanisms and Improvements

Dayal Singh Kalra, Maissam Barkeshli

2024年份
87被引次数
9顶会引用

摘要

It is common in deep learning to warm up the learning rate η\eta, often by a linear schedule between ηinit=0\eta_{\text{init}} = 0 and a predetermined target ηtrgt\eta_{\text{trgt}}. In this paper, we show through systematic experiments using SGD and Adam that the overwhelming benefit of warmup arises from allowing the network to tolerate larger ηtrgt\eta_{\text{trgt}} by forcing the network to more well-conditioned areas of the loss landscape. The ability to handle larger ηtrgt\eta_{\text{trgt}} makes hyperparameter tuning more robust while improving the final performance. We uncover different regimes of operation during the warmup period, depending on whether training starts off in a progressive sharpening or sharpness reduction phase, which in turn depends on the initialization and parameterization. Using these insights, we show how ηinit\eta_{\text{init}} can be properly chosen by utilizing the loss catapult mechanism, which saves on the number of warmup steps, in some cases completely eliminating the need for warmup. We also suggest an initialization for the variance in Adam which provides benefits similar to warmup.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4bcb17dd-a029-45fe-89cb-2b7c91fca9e2

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖