A Dynamical Central Limit Theorem for Shallow Neural Networks
Zhengdao Chen, Grant M. Rotskoff, Joan Bruna, Eric Vanden-Eijnden
Abstract
Recent theoretical work has characterized the dynamics of wide shallow neural networks trained via gradient descent in an asymptotic regime called the mean-field limit as the number of parameters tends towards infinity. At initialization, the randomly sampled parameters lead to a deviation from the mean-field limit that is dictated by the classical Central Limit Theorem (CLT). However, the dynamics of training introduces correlations among the parameters, raising the question of how the fluctuations evolve during training. Here, we analyze the mean-field dynamics as a Wasserstein gradient flow and prove that the deviations from the mean-field limit scaled by the width, in the width-asymptotic limit, remain bounded throughout training. In particular, they eventually vanish in the CLT scaling if the mean-field dynamics converges to a measure that interpolates the training data. This observation has implications for both the approximation rate and the generalization: the upper bound we obtain is given by a Monte-Carlo type resampling error, which does not depend explicitly on the dimension. This bound motivates a regularizaton term on the 2-norm of the underlying measure, which is also connected to generalization via the variation-norm function spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81188ccd-0cde-4a64-bd3a-9a859b6554ffCited by top-tier papers7
- Embedding Principle of Loss Landscape of Deep Neural NetworksYaoyu Zhang, Zhongwang Zhang, Tao Luo, Zhi-Qin John XuNeurIPS 2021 · 48 citations
- Feature-Learning Networks Are Consistent Across Widths At Realistic ScalesNikhil Vyas, Alexander B. Atanasov, Blake Bordelon, Depen Morwani et al.NeurIPS 2023 · 47 citations
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 10 citations
- Limiting fluctuation and trajectorial stability of multilayer neural networks with mean field trainingHuy Tuan Pham, Phan-Minh NguyenNeurIPS 2021 · 6 citations
- Symmetries in Overparametrized Neural Networks: A Mean Field ViewJavier Maass Martínez, Joaquín FontbonaNeurIPS 2024 · 4 citations
Builds on7
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
Related papers
- Global Convergence of Three-layer Neural Networks in the Mean Field RegimeHuy Tuan Pham, Phan-Minh NguyenICLR 2021 · 23 citations
- Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2023 · 56 citations
- Quantitative convergence of trained neural networks to Gaussian processesAndrea Agazzi, Eloy Mósig García, Dario TrevisanNeurIPS 2025
- Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning SystemsOri Shem-Ur, Khen Cohen, Yaron OzICLR 2026
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos et al.ICLR 2024 · 2 citations
