Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three Regimes
Maxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. Vetrov
摘要
A fundamental property of deep learning normalization techniques, such as batch normalization, is making the pre-normalization parameters scale invariant. The intrinsic domain of such parameters is the unit sphere, and therefore their gradient optimization dynamics can be represented via spherical optimization with varying effective learning rate (ELR), which was studied previously. However, the varying ELR may obscure certain characteristics of the intrinsic loss landscape structure. In this work, we investigate the properties of training scale-invariant neural networks directly on the sphere using a fixed ELR. We discover three regimes of such training depending on the ELR value: convergence, chaotic equilibrium, and divergence. We study these regimes in detail both on a theoretical examination of a toy example and on a thorough empirical analysis of real scale-invariant deep learning models. Each regime has unique features and reflects specific properties of the intrinsic loss landscape, some of which have strong parallels with previous research on both regular and scale-invariant neural networks training. Finally, we demonstrate how the discovered regimes are reflected in conventional training of normalized networks and how they can be leveraged to achieve better optima.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 被引用 101 次
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens 等NeurIPS 2024 · 被引用 69 次
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- Towards Practical Control of Singular Values of Convolutional LayersAlexandra Senderovich, Ekaterina Bulatova, Anton Obukhov, Maxim V. RakhubaNeurIPS 2022 · 被引用 16 次
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva 等NeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper15
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 被引用 235 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
相关 Paper
- Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight DecayRuosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian SunNeurIPS 2021 · 被引用 46 次
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 被引用 93 次
- Fast Equilibrium of SGD in Generic SituationsZhiyuan Liu, Yi Wang, Zhiren WangICLR 2024 · 被引用 1 次
- New Interpretations of Normalization Methods in Deep LearningJiacheng Sun, Xiangyong Cao, Hanwen Liang, Weiran Huang 等AAAI 2020 · 被引用 39 次
- On skip connections and normalisation layers in deep optimisationLachlan E. MacDonald, Jack Valmadre, Hemanth Saratchandran, Simon LuceyNeurIPS 2023 · 被引用 8 次
