Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three Regimes
Maxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. Vetrov
Abstract
A fundamental property of deep learning normalization techniques, such as batch normalization, is making the pre-normalization parameters scale invariant. The intrinsic domain of such parameters is the unit sphere, and therefore their gradient optimization dynamics can be represented via spherical optimization with varying effective learning rate (ELR), which was studied previously. However, the varying ELR may obscure certain characteristics of the intrinsic loss landscape structure. In this work, we investigate the properties of training scale-invariant neural networks directly on the sphere using a fixed ELR. We discover three regimes of such training depending on the ELR value: convergence, chaotic equilibrium, and divergence. We study these regimes in detail both on a theoretical examination of a toy example and on a thorough empirical analysis of real scale-invariant deep learning models. Each regime has unique features and reflects specific properties of the intrinsic loss landscape, some of which have strong parallels with previous research on both regular and scale-invariant neural networks training. Finally, we demonstrate how the discovered regimes are reflected in conventional training of normalized networks and how they can be leveraged to achieve better optima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dd5fec4-872e-49e2-bc5a-c3de864c74e6Cited by top-tier papers7
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens et al.NeurIPS 2024 · 69 citations
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- Towards Practical Control of Singular Values of Convolutional LayersAlexandra Senderovich, Ekaterina Bulatova, Anton Obukhov, Maxim V. RakhubaNeurIPS 2022 · 16 citations
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva et al.NeurIPS 2024 · 6 citations
Builds on15
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
Related papers
- Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight DecayRuosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian SunNeurIPS 2021 · 46 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
- Fast Equilibrium of SGD in Generic SituationsZhiyuan Liu, Yi Wang, Zhiren WangICLR 2024 · 1 citation
- New Interpretations of Normalization Methods in Deep LearningJiacheng Sun, Xiangyong Cao, Hanwen Liang, Weiran Huang et al.AAAI 2020 · 39 citations
- On skip connections and normalisation layers in deep optimisationLachlan E. MacDonald, Jack Valmadre, Hemanth Saratchandran, Simon LuceyNeurIPS 2023 · 8 citations
