Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
Soham De, Samuel L. Smith
摘要
Batch normalization dramatically increases the largest trainable depth of residual networks, and this benefit has been crucial to the empirical success of deep residual networks on a wide range of benchmarks. We show that this key benefit arises because, at initialization, batch normalization downscales the residual branch relative to the skip connection, by a normalizing factor on the order of the square root of the network depth. This ensures that, early in training, the function computed by normalized residual blocks in deep networks is close to the identity function (on average). We use this insight to develop a simple initialization scheme that can train deep residual networks without normalization. We also provide a detailed empirical study of residual networks, which clarifies that, although batch normalized networks can be trained with larger learning rates, this effect is only beneficial in specific compute regimes, and has minimal benefits when the batch size is small.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 被引用 1,298 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
- GemNet: Universal Directional Graph Neural Networks for MoleculesJohannes Gasteiger, Florian Becker, Stephan GünnemannNeurIPS 2021 · 被引用 665 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
它引用的顶会 Paper4
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
相关 Paper
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann 等NeurIPS 2020 · 被引用 73 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 被引用 6 次
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 被引用 50 次
- Training Deep Spiking Neural Networks without NormalizationXinyu Shi, Zhaofei YuICML 2026
