Group Whitening: Balancing Learning Efficiency and Representational Capacity
Lei Huang, Yi Zhou, Li Liu, Fan Zhu, Ling Shao
Abstract
Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model's learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population statistics for inference can be avoided through group normalization (GN). This paper proposes group whitening (GW), which exploits the advantages of the whitening operation and avoids the disadvantages of normalization within mini-batches. In addition, we analyze the constraints imposed on features by normalization, and show how the batch size (group number) affects the performance of batch (group) normalized networks, from the perspective of model's representational capacity . This analysis provides theoretical guidance for applying GW in practice. Finally, we apply the proposed GW to ResNet and ResNeXt architectures and conduct experiments on the ImageNet and COCO benchmarks. Results show that GW consistently improves the performance of different architectures, with absolute gains of 1.02% ⇠ 1.49% in top-1 accuracy on ImageNet and 1.82% ⇠ 3.21% in bounding box AP on COCO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1017e6d-7e25-40e7-8f20-4198f1e2b8d1Cited by top-tier papers4
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo et al.CVPR 2022 · 25 citations
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceAntoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2021 · 22 citations
- Fast Differentiable Matrix Square RootYue Song, Nicu Sebe, Wei WangICLR 2022 · 18 citations
- On the Nonlinearity of Layer NormalizationYunhao Ni, Yuxin Guo, Junlong Jia, Lei HuangICML 2024 · 9 citations
Builds on7
- GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingTianle Cai, Shengjie Luo, Keyulu Xu, Di He et al.ICML 2021 · 224 citations
- Switchable Whitening for Deep Representation LearningXingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang et al.ICCV 2019 · 204 citations
- On the Number of Linear Regions of Convolutional Neural NetworksHuan Xiong, Lei Huang, Mengyang Yu, Li Liu et al.ICML 2020 · 80 citations
- EvalNorm: Estimating Batch Normalization Statistics for EvaluationSaurabh Singh, Abhinav ShrivastavaICCV 2019 · 57 citations
- Channel Equilibrium Networks for Learning Deep RepresentationWenqi Shao, Shitao Tang, Xingang Pan, Ping Tan et al.ICML 2020 · 17 citations
Related papers
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu et al.CVPR 2020
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Improving Generalization of Batch Whitening by Convolutional Unit OptimizationYooshin Cho, Hanbyel Cho, Youngsoo Kim, Junmo KimICCV 2021 · 3 citations
- Four Things Everyone Should Know to Improve Batch NormalizationCecilia Summers, Michael J. DinneenICLR 2020 · 57 citations
- Towards Stabilizing Batch Statistics in Backward Propagation of Batch NormalizationJunjie Yan, Ruosi Wan, Xiangyu Zhang, Wei Zhang et al.ICLR 2020 · 42 citations
