Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
Zhiqi Bu, Shiyun Xu, Jialin Mao
Abstract
Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work, we examine the applicability of convexity and Lipschitz continuity in deep learning, in order to precisely control the loss dynamics via the learning rate schedules. We illustrate that deep learning quickly becomes weakly convex after a short period of training, and the loss is predicable by an upper bound on the last iterate, which further informs the scaling of optimal learning rate. Through the lens of convexity, we build scaling laws of learning rates and losses that extrapolate as much as across training horizons and across model sizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf8ccbdc-85c5-43b9-bedd-5a4901503cd8Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal et al.NeurIPS 2024 · 168 citations
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
Related papers
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization PathsCharles Guille-Escuret, Hiroki Naganuma, Kilian Fatras, Ioannis MitliagkasICML 2024 · 9 citations
- Mechanic: A Learning Rate TunerAshok Cutkosky, Aaron Defazio, Harsh MehtaNeurIPS 2023 · 27 citations
- On the training dynamics of deep networks with regularizationAitor Lewkowycz, Guy Gur-AriNeurIPS 2020 · 27 citations
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 21 citations
- Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural NetworksShikai Qiu, Lechao Xiao, Andrew Gordon Wilson, Jeffrey Pennington et al.ICML 2025
