Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch Dependence
Antoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo Luschi
Abstract
We investigate the reasons for the performance degradation incurred with batch-independent normalization. We find that the prototypical techniques of layer normalization and instance normalization both induce the appearance of failure modes in the neural network’s pre-activations: (i) layer normalization induces a collapse towards channel-wise constant functions; (ii) instance normalization induces a lack of variability in instance statistics, symptomatic of an alteration of the expressivity. To alleviate failure mode (i) without aggravating failure mode (ii), we introduce the technique “Proxy Normalization” that normalizes post-activations using a proxy distribution. When combined with layer normalization or group normalization, this batch-independent normalization emulates batch normalization’s behavior and consistently matches or exceeds its performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1f867ef-94d8-4c70-811c-e419a5101330Cited by top-tier papers8
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 50 citations
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 19 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
- On the Nonlinearity of Layer NormalizationYunhao Ni, Yuxin Guo, Junlong Jia, Lei HuangICML 2024 · 9 citations
Builds on15
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Evolving Normalization-Activation LayersHanxiao Liu, Andy Brock, Karen Simonyan, Quoc LeNeurIPS 2020 · 94 citations
Related papers
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo et al.CVPR 2022 · 25 citations
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 6 citations
- Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer NormalizationMinKyu Lee, Sangeek Hyun, Woojin Jun, Hyunjun Kim et al.ICLR 2026
- CrossNorm and SelfNorm for Generalization under Distribution ShiftsZhiqiang Tang, Yunhe Gao, Yi Zhu, Zhi Zhang et al.ICCV 2021 · 70 citations
