On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight Decay
Ekaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin, Dmitry P. Vetrov
摘要
Training neural networks with batch normalization and weight decay has become a common practice in recent years. In this work, we show that their combined use may result in a surprising periodic behavior of optimization dynamics: the training process regularly exhibits destabilizations that, however, do not lead to complete divergence but cause a new period of training. We rigorously investigate the mechanism underlying the discovered periodic behavior from both empirical and theoretical points of view and analyze the conditions in which it occurs in practice. We also demonstrate that periodic behavior can be regarded as a generalization of two previously opposing perspectives on training with batch normalization and weight decay, namely the equilibrium presumption and the instability presumption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens 等NeurIPS 2024 · 被引用 69 次
- Adapting the Linearised Laplace Model Evidence for Modern Deep LearningJavier Antorán, David Janz, James Urquhart Allingham, Erik A. Daxberger 等ICML 2022 · 被引用 36 次
- Robust Training of Neural Networks Using Scale Invariant ArchitecturesZhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank J. Reddi 等ICML 2022 · 被引用 33 次
- Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three RegimesMaxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. VetrovNeurIPS 2022 · 被引用 25 次
它引用的顶会 Paper4
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 被引用 93 次
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter 等ICLR 2021 · 被引用 22 次
- Understanding the Disharmony between Weight Normalization Family and Weight DecayXiang Li, Shuo Chen, Jian YangAAAI 2020 · 被引用 18 次
相关 Paper
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- On the Training Instability of Shuffling SGD with Batch NormalizationDavid Xing Wu, Chulhee Yun, Suvrit SraICML 2023 · 被引用 6 次
- Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight DecayRuosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian SunNeurIPS 2021 · 被引用 46 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 被引用 101 次
