Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
Andrea Montanari, Pierfrancesco Urbani
摘要
Understanding the inductive bias and generalization properties of large overparametrized machine learning models requires to characterize the dynamics of the training algorithm. We study the learning dynamics of large two-layer neural networks via dynamical mean field theory, a well established technique of non-equilibrium statistical physics. We show that, for large network width , and large number of samples per input dimension , the training dynamics exhibits a separation of timescales which implies: The emergence of a slow time scale associated with the growth in Gaussian/Rademacher complexity of the network; Inductive bias towards small complexity if the initialization has small enough complexity; A dynamical decoupling between feature learning and overfitting regimes; A non-monotone behavior of the test error, associated `feature unlearning'regime at large times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 被引用 14 次
- The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic NetworksVittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 被引用 13 次
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala 等NeurIPS 2025 · 被引用 10 次
- Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic LimitBohan Zhang, Zihao Wang, Hengyu Fu, Jason D. LeeICLR 2026 · 被引用 3 次
- Biased Generalization in Diffusion ModelsLuca Saglietti, Luca Biggio, Jerome Garnier-Brun, Davide Beltrame 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper4
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 被引用 95 次
相关 Paper
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient DescentShota Imai, Sota Nishiyama, Masaaki ImaizumiICML 2026
- Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvaluesJakob Kramp, Javed Lindner, Moritz HeliasICML 2026 · 被引用 2 次
- An analytic theory of shallow networks dynamics for hinge loss classificationFranco Pellegrini, Giulio BiroliNeurIPS 2020 · 被引用 19 次
- Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2023 · 被引用 56 次
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 被引用 33 次
