Towards Training Without Depth Limits: Batch Normalization Without Gradient Explosion
Alexandru Meterez, Amir Joudaki, Francesco Orabona, Alexander Immer, Gunnar Rätsch, Hadi Daneshmand
摘要
Normalization layers are one of the key building blocks for deep neural networks. Several theoretical studies have shown that batch normalization improves the signal propagation, by avoiding the representations from becoming collinear across the layers. However, results on mean-field theory of batch normalization also conclude that this benefit comes at the expense of exploding gradients in depth. Motivated by these two aspects of batch normalization, in this study we pose the following question: "Can a batch-normalized network keep the optimal signal propagation properties, but avoid exploding gradients?" We answer this question in the affirmative by giving a particular construction of an Multi-Layer Perceptron (MLP) with linear activations and batch-normalization that provably has bounded gradients at any depth. Based on Weingarten calculus, we develop a rigorous and non-asymptotic theory for this constructed MLP that gives a precise characterization of forward signal propagation, while proving that gradients remain bounded for linearly independent input samples, which holds in most practical settings. Inspired by our theory, we also design an activation shaping scheme that empirically achieves the same properties for certain nonlinear activations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph LearningAlvaro Arroyo, Alessio Gravina, Benjamin Gutteridge, Federico Barbero 等NeurIPS 2025 · 被引用 58 次
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 被引用 52 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
- On the impact of activation and normalization in obtaining isometric embeddings at initializationAmir Joudaki, Hadi Daneshmand, Francis R. BachNeurIPS 2023 · 被引用 16 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 等NeurIPS 2022 · 被引用 161 次
相关 Paper
- On Bridging the Gap between Mean Field and Finite Width Deep Random Multilayer Perceptron with Batch NormalizationAmir Joudaki, Hadi Daneshmand, Francis R. BachICML 2023 · 被引用 4 次
- Characterizing signal propagation to close the performance gap in unnormalized ResNetsAndrew Brock, Soham De, Samuel L. SmithICLR 2021 · 被引用 21 次
- On skip connections and normalisation layers in deep optimisationLachlan E. MacDonald, Jack Valmadre, Hemanth Saratchandran, Simon LuceyNeurIPS 2023 · 被引用 8 次
- Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningEkdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2021 · 被引用 50 次
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 被引用 51 次
