Is normalization indispensable for training deep neural network?
Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue, Bhiksha Raj
Abstract
Normalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, one question can be whether normalization is indispensable for training deep neural networks? In this paper, we analyze what would happen when normalization layers are removed from the networks, and show how to train deep neural networks without normalization layers and without performance degradation. Our proposed method can achieve the same or even slightly better performance in a variety of tasks: image classification in ImageNet, object detection and segmentation in MS-COCO, video classification in Kinetics, and machine translation in WMT English-German, etc. Our study may help better understand the role of normalization layers and can be a competitive alternative to normalization layers. Codes are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4306d73-face-4afe-9c35-24c8791de249Cited by top-tier papers5
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- TokenScout: Early Detection of Ethereum Scam Tokens via Temporal Graph LearningCong Wu, Jing Chen, Ziming Zhao, Kun He et al.CCS 2024 · 35 citations
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 22 citations
- Catformer: Designing Stable Transformers via Sensitivity AnalysisJared Quincy Davis, Albert Gu, Krzysztof Choromanski, Tri Dao et al.ICML 2021 · 19 citations
- Wide Bayesian neural networks have a simple weight posterior: theory and accelerated samplingJiri Hron, Roman Novak, Jeffrey Pennington, Jascha Sohl-DicksteinICML 2022 · 10 citations
Builds on1
Related papers
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma et al.ICML 2020 · 57 citations
- Transformers without NormalizationJiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun et al.CVPR 2025
- Network DeconvolutionChengxi Ye, Matthew Evanusa, Hua He, Anton Mitrokhin et al.ICLR 2020
- TaskNorm: Rethinking Batch Normalization for Meta-LearningJohn Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin et al.ICML 2020 · 93 citations
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang et al.ICLR 2020 · 42 citations
