Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch Dependence
Antoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo Luschi
摘要
We investigate the reasons for the performance degradation incurred with batch-independent normalization. We find that the prototypical techniques of layer normalization and instance normalization both induce the appearance of failure modes in the neural network’s pre-activations: (i) layer normalization induces a collapse towards channel-wise constant functions; (ii) instance normalization induces a lack of variability in instance statistics, symptomatic of an alteration of the expressivity. To alleviate failure mode (i) without aggravating failure mode (ii), we introduce the technique “Proxy Normalization” that normalizes post-activations using a proxy distribution. When combined with layer normalization or group normalization, this batch-independent normalization emulates batch normalization’s behavior and consistently matches or exceeds its performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 被引用 50 次
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 被引用 19 次
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung 等ICML 2024 · 被引用 16 次
- On the Nonlinearity of Layer NormalizationYunhao Ni, Yuxin Guo, Junlong Jia, Lei HuangICML 2024 · 被引用 9 次
它引用的顶会 Paper15
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 被引用 173 次
- Evolving Normalization-Activation LayersHanxiao Liu, Andy Brock, Karen Simonyan, Quoc LeNeurIPS 2020 · 被引用 94 次
相关 Paper
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo 等CVPR 2022 · 被引用 25 次
- Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural NetworksSaurabh Singh, Shankar KrishnanCVPR 2020
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 被引用 6 次
- Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer NormalizationMinKyu Lee, Sangeek Hyun, Woojin Jun, Hyunjun Kim 等ICLR 2026
- CrossNorm and SelfNorm for Generalization under Distribution ShiftsZhiqiang Tang, Yunhe Gao, Yi Zhu, Zhi Zhang 等ICCV 2021 · 被引用 70 次
