On Bridging the Gap between Mean Field and Finite Width Deep Random Multilayer Perceptron with Batch Normalization
Amir Joudaki, Hadi Daneshmand, Francis R. Bach
Abstract
Mean-field theory is widely used in theoretical studies of neural networks. In this paper, we analyze the role of depth in the concentration of mean-field predictions for Gram matrices of hidden representations in deep multilayer perceptron (MLP) with batch normalization (BN) at initialization. It is postulated that the mean-field predictions suffer from layer-wise errors that amplify with depth. We demonstrate that BN avoids this error amplification with depth. When the chain of hidden representations is rapidly mixing, we establish a concentration bound for a mean-field model of Gram matrices. To our knowledge, this is the first concentration bound that does not become vacuous with depth for standard MLPs with a finite width.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b59b6507-627e-42ed-8fb5-b9325d0a76a1Cited by top-tier papers3
- Towards Training Without Depth Limits: Batch Normalization Without Gradient ExplosionAlexandru Meterez, Amir Joudaki, Francesco Orabona, Alexander Immer et al.ICLR 2024 · 11 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- On the Importance of Gaussianizing RepresentationsDaniel Eftekhari, Vardan PapyanICML 2025
Builds on9
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu et al.ICLR 2020 · 77 citations
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann et al.NeurIPS 2020 · 73 citations
Related papers
- On the impact of activation and normalization in obtaining isometric embeddings at initializationAmir Joudaki, Hadi Daneshmand, Francis R. BachNeurIPS 2023 · 16 citations
- Batch Normalization Orthogonalizes Representations in Deep Random NetworksHadi Daneshmand, Amir Joudaki, Francis R. BachNeurIPS 2021 · 47 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
- Limiting fluctuation and trajectorial stability of multilayer neural networks with mean field trainingHuy Tuan Pham, Phan-Minh NguyenNeurIPS 2021 · 6 citations
- Theoretical Characterisation of the Gauss Newton Conditioning in Neural NetworksJim Zhao, Sidak Pal Singh, Aurélien LucchiNeurIPS 2024 · 7 citations
