On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight Decay
Ekaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin, Dmitry P. Vetrov
Abstract
Training neural networks with batch normalization and weight decay has become a common practice in recent years. In this work, we show that their combined use may result in a surprising periodic behavior of optimization dynamics: the training process regularly exhibits destabilizations that, however, do not lead to complete divergence but cause a new period of training. We rigorously investigate the mechanism underlying the discovered periodic behavior from both empirical and theoretical points of view and analyze the conditions in which it occurs in practice. We also demonstrate that periodic behavior can be regarded as a generalization of two previously opposing perspectives on training with batch normalization and weight decay, namely the equilibrium presumption and the instability presumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55174f6a-5edf-4765-bf20-17f3bf900fa2Cited by top-tier papers10
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens et al.NeurIPS 2024 · 69 citations
- Adapting the Linearised Laplace Model Evidence for Modern Deep LearningJavier Antorán, David Janz, James Urquhart Allingham, Erik A. Daxberger et al.ICML 2022 · 36 citations
- Robust Training of Neural Networks Using Scale Invariant ArchitecturesZhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank J. Reddi et al.ICML 2022 · 33 citations
- Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three RegimesMaxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. VetrovNeurIPS 2022 · 25 citations
Builds on4
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter et al.ICLR 2021 · 22 citations
- Understanding the Disharmony between Weight Normalization Family and Weight DecayXiang Li, Shuo Chen, Jian YangAAAI 2020 · 18 citations
Related papers
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- On the Training Instability of Shuffling SGD with Batch NormalizationDavid Xing Wu, Chulhee Yun, Suvrit SraICML 2023 · 6 citations
- Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight DecayRuosi Wan, Zhanxing Zhu, Xiangyu Zhang, Jian SunNeurIPS 2021 · 46 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
