Quantitative convergence of trained neural networks to Gaussian processes
Andrea Agazzi, Eloy Mósig García, Dario Trevisan
摘要
In this paper, we study the quantitative convergence of shallow neural networks trained via gradient descent to their associated Gaussian processes in the infinitewidth limit. While previous work has established qualitative convergence under broad settings, precise, finite-width estimates remain limited, particularly during training. We provide explicit upper bounds on the quadratic Wasserstein distance between the network output and its Gaussian approximation at any training time t ≥ 0, demonstrating polynomial decay with network width. Our results quantify how architectural parameters, such as width and input dimension, influence convergence, and how training dynamics affect the approximation error. However, these results were largely confined to the initialization regime. To this day, extensions to the full training trajectory remained limited, with few works addressing how approximation errors evolve over time or depend on architectural features such as width and depth. The present work builds on this gap by extending the quantitative convergence discussed above to trained networks, providing explicit bounds on the Wasserstein distance between the network output and the associated Gaussian process for any positive training time. From a spectral perspective, the NTK's conditioning plays a central role in understanding convergence rates and generalization. Lower bounds on the smallest eigenvalue of the empirical NTK have been derived under various conditions. For instance, Karhadkar et al. [2024] and Bombari et al. [2022] provide sharp bounds in the context of ReLU and smooth activation functions, respectively. Additionaly, Carvalho et al. [2025] showed that under very mild assumptions on the non-linearity and non-proportionality of the training data, the analytic NTK is not degenerate. These results are essential for establishing the stability of the gradient flow and, hence, for deriving quantitative convergence guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
- Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationSimone Bombari, Mohammad Hossein Amani, Marco MondelliNeurIPS 2022 · 被引用 45 次
- Smooth p-Wasserstein Distance: Structure, Empirical Approximation, and Statistical ApplicationsSloan Nietert, Ziv Goldfeld, Kengo KatoICML 2021 · 被引用 39 次
相关 Paper
- A Dynamical Central Limit Theorem for Shallow Neural NetworksZhengdao Chen, Grant M. Rotskoff, Joan Bruna, Eric Vanden-EijndenNeurIPS 2020 · 被引用 33 次
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
- Spectral Bias Outside the Training Set for Deep Networks in the Kernel RegimeBenjamin Bowman, Guido F. MontúfarNeurIPS 2022 · 被引用 17 次
- Explicit loss asymptotics in the gradient descent training of neural networksMaksim Velikanov, Dmitry YarotskyNeurIPS 2021 · 被引用 19 次
- Characterizing the spectrum of the NTK via a power series expansionMichael Murray, Hui Jin, Benjamin Bowman, Guido MontúfarICLR 2023 · 被引用 2 次
