Lune

ICLR2025顶会

Taming Transformer Without Using Learning Rate Warmup

Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou, Rong Xiao

出版方
2025年份
4顶会引用

摘要

Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and using an obviously lower learning rate is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal the rationale behind the model crash phenomenon in the training process, termed spectral energy concentration of W q ⊤ W k , which is the reason for a malignant entropy collapse, where W q and W k are the projection matrices for the query and the key in Transformer, respectively. To remedy this problem, motivated by Weyl's Inequality, we present a novel optimization strategy, i.e., making the weight updating in successive steps smooth-if the ratio ) is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of σ 1 (W t -1 ) σ 1 (∇W t ) , where ∇W t is the updating quantity in step t . Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these Transformers without using learning rate warmup.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖