Towards a General Theory of Infinite-Width Limits of Neural Classifiers
Eugene A. Golikov
摘要
Obtaining theoretical guarantees for neural networks training appears to be a hard problem in a general case. Recent research has been focused on studying this problem in the limit of infinite width and two different theories have been developed: a mean-field (MF) and a constant kernel (NTK) limit theories. We propose a general framework that provides a link between these seemingly distinct theories. Our framework out of the box gives rise to a discrete-time MF limit which was not previously explored in the literature. We prove a convergence theorem for it, and show that it provides a more reasonable approximation for finite-width nets compared to the NTK limit if learning rates are not very small. Also, our framework suggests a limit model that coincides neither with the MF limit nor with the NTK one. We show that for networks with more than two hidden layers RMSProp training has a non-trivial discrete-time MF limit but GD training does not have one. Overall, our framework demonstrates that both MF and NTK limits have considerable limitations in approximating finite-sized neural nets, indicating the need for designing more accurate infinite-width approximations for them.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process PerspectiveGeoff Pleiss, John P. CunninghamNeurIPS 2021 · 被引用 35 次
- Explicit loss asymptotics in the gradient descent training of neural networksMaksim Velikanov, Dmitry YarotskyNeurIPS 2021 · 被引用 19 次
- Non-Gaussian Tensor ProgramsEugene A. Golikov, Greg YangNeurIPS 2022 · 被引用 11 次
- Multi-Layer Neural Networks as Trainable Ladders of Hilbert SpacesZhengdao ChenICML 2023 · 被引用 4 次
- Gradient Flow Through Diagram Expansions: Learning Regimes and Explicit SolutionsDmitry Yarotsky, Eugene Golikov, Yaroslav GusevICML 2026 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width AnalysisHoang Pham, The Anh Ta, Tom Jacobs, Rebekka Burkholz 等NeurIPS 2025 · 被引用 2 次
- Adaptive Optimization in the ∞-Width LimitEtai Littwin, Greg YangICLR 2023
- Global Convergence of Three-layer Neural Networks in the Mean Field RegimeHuy Tuan Pham, Phan-Minh NguyenICLR 2021 · 被引用 23 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
- A generalized neural tangent kernel for surrogate gradient learningLuke Eilers, Raoul-Martin Memmesheimer, Sven GoedekeNeurIPS 2024 · 被引用 2 次
