Deconstructing the Regularization of BatchNorm
Yann N. Dauphin, Ekin Dogus Cubuk
Abstract
Batch normalization (BatchNorm) has become a standard technique in deep learning. Its popularity is in no small part due to its often positive effect on generalization. Despite this success, the regularization effect of the technique is still poorly understood. This study aims to decompose BatchNorm into separate mechanisms that are much simpler. We identify three effects of BatchNorm and assess their impact directly with ablations and interventions. Our experiments show that preventing explosive growth at the final layer at initialization and during training can recover a large part of BatchNorm's generalization boost. This regularization mechanism can lift accuracy by 2.9% for Resnet-50 on Imagenet without BatchNorm. We show it is linked to other methods like Dropout and recent initializations like Fixup. Surprisingly, this simple mechanism matches the improvement of 0.9% of the more complex Dropout regularization for the state-of-the-art Efficientnet-B8 model on Imagenet. This demonstrates the underrated effectiveness of simple regularizations and sheds light on directions to further improve generalization for deep nets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3df57f1-af9f-4a98-89fb-1d7cb33bd94eCited by top-tier papers6
- Why Do Better Loss Functions Lead to Less Transferable Features?Simon Kornblith, Ting Chen, Honglak Lee, Mohammad NorouziNeurIPS 2021 · 113 citations
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 22 citations
- Feature Kernel DistillationBobby He, Mete OzayICLR 2022 · 18 citations
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan et al.NeurIPS 2025 · 5 citations
Builds on3
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Adversarial Examples Improve Image RecognitionCihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang et al.CVPR 2020
- SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationXianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi et al.CVPR 2020
Related papers
- Channel Regeneration: Improving Channel Utilization for Compact DNNsAnkit Kumar Sharma, Hassan ForooshAAAI 2023
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Characterizing signal propagation to close the performance gap in unnormalized ResNetsAndrew Brock, Soham De, Samuel L. SmithICLR 2021 · 21 citations
- Instance Enhancement Batch Normalization: An Adaptive Regulator of Batch NoiseSenwei Liang, Zhongzhan Huang, Mingfu Liang, Haizhao YangAAAI 2020 · 65 citations
