Delving into the Estimation Shift of Batch Normalization in a Network
Lei Huang, Yi Zhou, Tian Wang, Jie Luo, Xianglong Liu
摘要
Batch normalization (BN) is a milestone technique in deep learning . It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation shift magnitude of BN to quantitatively measure the difference between its estimated population statistics and expected ones. Our primary observation is that the estimation shift can be accumulated due to the stack of BN in a network, which has detriment effects for the test performance. We further find a batch-free normalization (BFN) can block such an accumulation of estimation shift. These observations motivate our design of XBNBlock that replace one BN with BFN in the bottleneck block of residual-style networks. Experiments on the ImageNet and COCO benchmarks show that XBNBlock consistently improves the performance of different architectures, including ResNet and ResNeXt, by a significant margin and seems to be more robust to distribution shift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Understanding the Failure of Batch Normalization for Transformers in NLPJiaxi Wang, Ji Wu, Lei HuangNeurIPS 2022 · 被引用 14 次
- Reducing Divergence in Batch Normalization for Domain AdaptationEllen Yi-Ge, Mingjing Wu, Zhenghan ChenAAAI 2025 · 被引用 3 次
它引用的顶会 Paper12
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann 等NeurIPS 2020 · 被引用 688 次
- U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image TranslationJunho Kim, Minjae Kim, Hyeonwoo Kang, Kwanghee LeeICLR 2020 · 被引用 632 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang 等ICCV 2019 · 被引用 204 次
- TaskNorm: Rethinking Batch Normalization for Meta-LearningJohn Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin 等ICML 2020 · 被引用 93 次
相关 Paper
- Group Whitening: Balancing Learning Efficiency and Representational CapacityLei Huang, Yi Zhou, Li Liu, Fan Zhu 等CVPR 2021
- EvalNorm: Estimating Batch Normalization Statistics for EvaluationSaurabh Singh, Abhinav ShrivastavaICCV 2019 · 被引用 57 次
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu 等CVPR 2020
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Removing Batch Normalization Boosts Adversarial TrainingHaotao Wang, Aston Zhang, Shuai Zheng, Xingjian Shi 等ICML 2022 · 被引用 51 次
