Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural Networks
Saurabh Singh, Shankar Krishnan
Abstract
Batch Normalization (BN) uses mini-batch statistics to normalize the activations during training, introducing dependence between mini-batch elements. This dependency can hurt the performance if the mini-batch size is too small, or if the elements are correlated. Several alternatives, such as Batch Renormalization and Group Normalization (GN), have been proposed to address this issue. However, they either do not match the performance of BN for large batches, or still exhibit degradation in performance for smaller batches, or introduce artificial constraints on the model architecture. In this paper we propose the Filter Response Normalization (FRN) layer, a novel combination of a normalization and an activation function, that can be used as a replacement for other normalizations and activations. Our method operates on each activation channel of each batch element independently, eliminating the dependency on other batch elements.
Our method outperforms BN and other alternatives in a variety of settings for all batch sizes. FRN layer performs « 0.7-1.0% better than BN on top-1 validation accuracy with large mini-batch sizes for Imagenet classification using InceptionV3 and ResnetV2-50 architectures. Further, it performs ą 1% better than GN on the same problem in the small mini-batch size regime. For object detection problem on COCO dataset, FRN layer outperforms all other methods by at least 0.3-0.5% in all batch size regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c8dcb3c-72a0-4d53-8f7d-0db312567802Cited by top-tier papers32
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Rethinking the Truly Unsupervised Image-to-Image TranslationKyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo et al.ICCV 2021 · 115 citations
- HyNet: Learning Local Descriptor with Hybrid Similarity Measure and Triplet LossYurun Tian, Axel Barroso Laguna, Tony Ng, Vassileios Balntas et al.NeurIPS 2020 · 101 citations
- Evolving Normalization-Activation LayersHanxiao Liu, Andy Brock, Karen Simonyan, Quoc LeNeurIPS 2020 · 94 citations
Builds on2
Related papers
- Group Whitening: Balancing Learning Efficiency and Representational CapacityLei Huang, Yi Zhou, Li Liu, Fan Zhu et al.CVPR 2021
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo et al.CVPR 2022 · 25 citations
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang et al.ICLR 2020 · 42 citations
- Local Context Normalization: Revisiting Local NormalizationAnthony Ortiz, Caleb Robinson, Dan Morris, Olac Fuentes et al.CVPR 2020
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
