Stochastic Whitening Batch Normalization
Shengdong Zhang, Ehsan Nezhadarya, Homa Fashandi, Jiayi Liu, Darin Graham, Mohak Shah
Abstract
Batch Normalization (BN) is a popular technique for training Deep Neural Networks (DNNs) . BN uses scaling and shifting to normalize activations of mini-batches to accelerate convergence and improve generalization. The recently proposed Iterative Normalization (IterNorm) method improves these properties by whitening the activations iteratively using Newton's method. However, since Newton's method initializes the whitening matrix independently at each training step, no information is shared between consecutive steps. In this work, instead of exact computation of whitening matrix at each time step, we estimate it gradually during training in an online fashion, using our proposed Stochastic Whitening Batch Normalization (SWBN) algorithm. We show that while SWBN improves the convergence rate and generalization of DNNs, its computational overhead is less than that of IterNorm. Due to the high efficiency of the proposed method, it can be easily employed in most DNN architectures with a large number of layers. We provide comprehensive experiments and comparisons between BN, IterNorm, and SWBN layers to demonstrate the effectiveness of the proposed technique in conventional (many-shot) image classification and few-shot classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Improving Generalization of Batch Whitening by Convolutional Unit OptimizationYooshin Cho, Hanbyel Cho, Youngsoo Kim, Junmo KimICCV 2021 · 3 citations
- An Investigation Into the Stochasticity of Batch WhiteningLei Huang, Lei Zhao, Yi Zhou, Fan Zhu et al.CVPR 2020
- Cross-Iteration Batch NormalizationZhuliang Yao, Yue Cao, Shuxin Zheng, Gao Huang et al.CVPR 2021
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan et al.CVPR 2020
- Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and AdaptationYilin Lyu, Liyuan Wang, Xingxing Zhang, Zicheng Sun et al.NeurIPS 2023 · 17 citations
