Maximal Initial Learning Rates in Deep ReLU Networks
Gaurav Iyer, Boris Hanin, David Rolnick
Abstract
Training a neural network requires choosing a suitable learning rate, which involves a trade-off between speed and effectiveness of convergence. While there has been considerable theoretical and empirical analysis of how large the learning rate can be, most prior work focuses only on late-stage training. In this work, we introduce the maximal initial learning rate - the largest learning rate at which a randomly initialized neural network can successfully begin training and achieve (at least) a given threshold accuracy. Using a simple approach to estimate , we observe that in constant-width fully-connected ReLU networks, behaves differently from the maximum learning rate later in training. Specifically, we find that is well predicted as a power of depth width, provided that (i) the width of the network is sufficiently large compared to the depth, and (ii) the input layer is trained at a relatively small learning rate. We further analyze the relationship between and the sharpness of the network at initialization, indicating they are closely though not inversely related. We formally prove bounds for in terms of depth width that align with our empirical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43017f3d-0da6-4740-a654-2dc50a9367ecCited by top-tier papers5
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 25 citations
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 21 citations
- On the Parameterization of Second-Order Optimization Effective towards the Infinite WidthSatoki Ishikawa, Ryo KarakidaICLR 2024 · 10 citations
- Principled Architecture-aware Scaling of HyperparametersWuyang Chen, Junru Wu, Zhangyang Wang, Boris HaninICLR 2024 · 3 citations
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial ConditionsGül Sena Altintas, Devin Kwok, Colin Raffel, David RolnickICML 2025
Builds on13
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
Related papers
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva et al.NeurIPS 2024 · 6 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy TrainingApostolos Evangelidis, Felix KrahmerICML 2026
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari et al.NeurIPS 2021 · 35 citations
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 87 citations
