Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep Learning
Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka
Abstract
Inspired by BatchNorm, there has been an explosion of normalization layers in deep learning. Recent works have identified a multitude of beneficial properties in BatchNorm to explain its success. However, given the pursuit of alternative normalization layers, these properties need to be generalized so that any given layer's success/failure can be accurately predicted. In this work, we take a first step towards this goal by extending known properties of BatchNorm in randomly initialized deep neural networks (DNNs) to several recently proposed normalization layers. Our primary findings follow: (i) similar to BatchNorm, activations-based normalization layers can prevent exponential growth of activations in ResNets, but parametric techniques require explicit remedies; (ii) use of GroupNorm can ensure an informative forward propagation, with different samples being assigned dissimilar activations, but increasing group size results in increasingly indistinguishable activations for different samples, explaining slow convergence speed in models with LayerNorm; and (iii) small group sizes result in large gradient norm in earlier layers, hence explaining training instability issues in Instance Normalization and illustrating a speed-stability tradeoff in GroupNorm. Overall, our analysis reveals a unified set of mechanisms that underpin the success of normalization methods in deep learning, providing us with a compass to systematically explore the vast design space of DNN normalization layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89de0192-2a08-40d8-9a32-20c1b58297cbCited by top-tier papers11
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 19 citations
- Stronger Normalization-Free TransformersMingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun et al.CVPR 2026 · 16 citations
- On the impact of activation and normalization in obtaining isometric embeddings at initializationAmir Joudaki, Hadi Daneshmand, Francis R. BachNeurIPS 2023 · 16 citations
- Magnitude Invariant Parametrizations Improve Hypernetwork LearningJose Javier Gonzalez Ortiz, John V. Guttag, Adrian V. DalcaICLR 2024 · 13 citations
Builds on20
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 672 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Neural Architecture Search without TrainingJoe Mellor, Jack Turner, Amos Storkey, Elliot J. CrowleyICML 2021 · 477 citations
Related papers
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 22 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 6 citations
- New Interpretations of Normalization Methods in Deep LearningJiacheng Sun, Xiangyong Cao, Hanwen Liang, Weiran Huang et al.AAAI 2020 · 39 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
