Weight Decay may matter more than µP for Learning Rate Transfer in Practice
Atli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi, Xi Chen
摘要
Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive. To this end, the Maximal Update Parameterization (µP) proposes a learning rate scaling designed to keep the update dynamics of internal representations stable across different model widths. However, the scaling rules of µP rely on strong assumptions, particularly about the geometric alignment of a layer’s inputs with both its weights and gradient updates. In this large-scale empirical investigation, we show that these assumptions hold only briefly at the start of training in the practical setups where learning rate transfer is most valuable, such as LLM training. For the remainder of training it is weight decay rather than µP that correctly stabilizes the update dynamics of internal representations across widths, facilitating learning rate transfer. This suggests µP's scaling primarily acts as a form of implicit learning rate warmup, allowing us to largely replace it with modified warmup schedules. Together these findings fundamentally challenge prevailing beliefs about learning rate transfer and can explain empirical observations such as why µP requires the independent weight decay variant for good transfer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
相关 Paper
- Scaling Exponents Across Parameterizations and OptimizersKatie E. Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi 等ICML 2024 · 被引用 59 次
- Learning Rate Scaling across LoRA Ranks and Transfer to Full FinetuningNan Chen, Soledad Villar, Soufiane HayouICML 2026 · 被引用 8 次
- Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural NetworksHaosong Zhang, Shenxi Wu, Xingjian Ma, Shirui Bian 等ICML 2026 · 被引用 1 次
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 被引用 8 次
