Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight Decay
Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian Sun
摘要
In this paper, we comprehensively reveal the learning dynamics of normalized neural network using Stochastic Gradient Descent (with momentum) and Weight Decay (WD), named as Spherical Motion Dynamics (SMD). Most related works on this topic focus on studying "effective learning rate" using "equilibrium" assumption, i.e. assuming weight norm has converge to a fixed value. However, their discussion on why equilibrium can be reached is either absent or unjustified. To clarify the mechanism behind, our work directly explores the cause of equilibrium, which should be regarded as a special state of SMD. Specifically, 1) we introduce the assumptions that can lead to equilibrium state in SMD, and prove equilibrium can be reached in a linear rate regime; 2) we propose "angular update" as a substitute for effective learning rate to depict the state of SMD, and derive the theoretical value of angular update in equilibrium state; 3) we verify our assumptions and theoretical results on various large-scale computer vision tasks including ImageNet and MSCOCO with standard settings. Experiment results show our theoretical findings agree well with empirical observations. Furthermore, we provide intuitive interpretations, showing how the behavior of angular update in SMD affects the optimization of neural network, and yields unexpected phenomenon in practice. We believe our findings and theoretical results can deepen our understanding on current training techniques for deep neural network.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Revisiting Weighted Aggregation in Federated Learning with Neural NetworksZexi Li, Tao Lin, Xinyi Shang, Chao WuICML 2023 · 被引用 119 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens 等NeurIPS 2024 · 被引用 69 次
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
它引用的顶会 Paper4
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins 等ICLR 2021 · 被引用 100 次
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 被引用 93 次
相关 Paper
- Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three RegimesMaxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. VetrovNeurIPS 2022 · 被引用 25 次
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight DecayEkaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin 等NeurIPS 2021 · 被引用 30 次
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural NetworksHidenori Tanaka, Daniel KuninNeurIPS 2021 · 被引用 57 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
