On the Weight Dynamics of Deep Normalized Networks
Christian H. X. Ali Mehmeti-Göpel, Michael Wand
Abstract
Recent studies have shown that high disparities in effective learning rates (ELRs) across layers in deep neural networks can negatively affect trainability. We formalize how these disparities evolve over time by modeling weight dynamics (evolution of expected gradient and weight norms) of networks with normalization layers, predicting the evolution of layer-wise ELR ratios. We prove that when training with any constant learning rate, ELR ratios converge to 1, despite initial gradient explosion. We identify a ``critical learning rate"beyond which ELR disparities widen, which only depends on current ELRs. To validate our findings, we devise a hyper-parameter-free warm-up method that successfully minimizes ELR spread quickly in theory and practice. Our experiments link ELR spread with trainability, a relationship that is most evident in very deep networks with significant gradient magnitude excursions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a706576-f33e-4211-acb6-bd79e9ac0e65Cited by top-tier papers3
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi et al.ICLR 2026 · 11 citations
- Layer-wise Update Aggregation with Recycling for Communication-Efficient Federated LearningJisoo Kim, Sungmin Kang, Sunwoo LeeNeurIPS 2025 · 4 citations
- Neural Differentiation in Deep Networks: A Theoretical Framework for Expressivity and Representational DiversityBoyuan Wang, Richard JiangCVPR 2026
Builds on8
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
Related papers
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- A Loss Curvature Perspective on Training Instabilities of Deep Learning ModelsJustin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta et al.ICLR 2022 · 49 citations
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
