Towards Stabilizing Batch Statistics in Backward Propagation of Batch Normalization
Junjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang, Yichen Wei, Jian Sun
Abstract
Batch Normalization (BN) is one of the most widely used techniques in Deep Learning field. But its performance can awfully degrade with insufficient batch size. This weakness limits the usage of BN on many computer vision tasks like detection or segmentation, where batch size is usually small due to the constraint of memory consumption. Therefore many modified normalization techniques have been proposed, which either fail to restore the performance of BN completely, or have to introduce additional nonlinear operations in inference procedure and increase huge consumption. In this paper, we reveal that there are two extra batch statistics involved in backward propagation of BN, on which has never been well discussed before. The extra batch statistics associated with gradients also can severely affect the training of deep neural network. Based on our analysis, we propose a novel normalization method, named Moving Average Batch Normalization (MABN). MABN can completely restore the performance of vanilla BN in small batch cases, without introducing any additional nonlinear operations in inference procedure. We prove the benefits of MABN by both theoretical analysis and experiments. Our experiments demonstrate the effectiveness of MABN in multiple computer vision tasks including ImageNet and COCO. The code has been released in https://github.com/megvii-model/MABN . * Equal Contribution. Work was done when Junjie Yan was an intern at Megvii Technology. † Corresponding author. 1 In the context of this paper, we use "batch size/normalization batch size" to refer the number of samples used to compute statistics unless otherwise stated. We use "gradient batch size" to refer the number of samples used to update weights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 630f0ab8-7ade-4d56-a9ac-5509cd51f007Cited by top-tier papers14
- GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingTianle Cai, Shengjie Luo, Keyulu Xu, Di He et al.ICML 2021 · 224 citations
- PowerNorm: Rethinking Batch Normalization in TransformersSheng Shen, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICML 2020 · 88 citations
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 40 citations
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo et al.CVPR 2022 · 25 citations
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 22 citations
Builds on2
Related papers
- Four Things Everyone Should Know to Improve Batch NormalizationCecilia Summers, Michael J. DinneenICLR 2020 · 57 citations
- Cross-Iteration Batch NormalizationZhuliang Yao, Yue Cao, Shuxin Zheng, Gao Huang et al.CVPR 2021
- Group Whitening: Balancing Learning Efficiency and Representational CapacityLei Huang, Yi Zhou, Li Liu, Fan Zhu et al.CVPR 2021
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Stochastic NormalizationZhi Kou, Kaichao You, Mingsheng Long, Jianmin WangNeurIPS 2020 · 138 citations
