Is normalization indispensable for training deep neural network?
Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue, Bhiksha Raj
摘要
Normalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, one question can be whether normalization is indispensable for training deep neural networks? In this paper, we analyze what would happen when normalization layers are removed from the networks, and show how to train deep neural networks without normalization layers and without performance degradation. Our proposed method can achieve the same or even slightly better performance in a variety of tasks: image classification in ImageNet, object detection and segmentation in MS-COCO, video classification in Kinetics, and machine translation in WMT English-German, etc. Our study may help better understand the role of normalization layers and can be a competitive alternative to normalization layers. Codes are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- TokenScout: Early Detection of Ethereum Scam Tokens via Temporal Graph LearningCong Wu, Jing Chen, Ziming Zhao, Kun He 等CCS 2024 · 被引用 35 次
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 被引用 22 次
- Catformer: Designing Stable Transformers via Sensitivity AnalysisJared Quincy Davis, Albert Gu, Krzysztof Choromanski, Tri Dao 等ICML 2021 · 被引用 19 次
- Wide Bayesian neural networks have a simple weight posterior: theory and accelerated samplingJiri Hron, Roman Novak, Jeffrey Pennington, Jascha Sohl-DicksteinICML 2022 · 被引用 10 次
它引用的顶会 Paper1
相关 Paper
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma 等ICML 2020 · 被引用 57 次
- Transformers without NormalizationJiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun 等CVPR 2025
- Network DeconvolutionChengxi Ye, Matthew Evanusa, Hua He, Anton Mitrokhin 等ICLR 2020
- TaskNorm: Rethinking Batch Normalization for Meta-LearningJohn Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin 等ICML 2020 · 被引用 93 次
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang 等ICLR 2020 · 被引用 42 次
