Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
Zhiqi Bu, Shiyun Xu, Jialin Mao
摘要
Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work, we examine the applicability of convexity and Lipschitz continuity in deep learning, in order to precisely control the loss dynamics via the learning rate schedules. We illustrate that deep learning quickly becomes weakly convex after a short period of training, and the loss is predicable by an upper bound on the last iterate, which further informs the scaling of optimal learning rate. Through the lens of convexity, we build scaling laws of learning rates and losses that extrapolate as much as across training horizons and across model sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal 等NeurIPS 2024 · 被引用 168 次
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li 等NeurIPS 2024 · 被引用 149 次
相关 Paper
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization PathsCharles Guille-Escuret, Hiroki Naganuma, Kilian Fatras, Ioannis MitliagkasICML 2024 · 被引用 9 次
- Mechanic: A Learning Rate TunerAshok Cutkosky, Aaron Defazio, Harsh MehtaNeurIPS 2023 · 被引用 27 次
- On the training dynamics of deep networks with regularizationAitor Lewkowycz, Guy Gur-AriNeurIPS 2020 · 被引用 27 次
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 被引用 21 次
- Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural NetworksShikai Qiu, Lechao Xiao, Andrew Gordon Wilson, Jeffrey Pennington 等ICML 2025
