An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini, Giulio Biroli
摘要
Neural networks have been shown to perform incredibly well in classification tasks over structured high-dimensional datasets. However, the learning dynamics of such networks is still poorly understood. In this paper we study in detail the training dynamics of a simple type of neural network: a single hidden layer trained to perform a classification task. We show that in a suitable mean-field limit this case maps to a single-node learning problem with a time-dependent dataset determined self-consistently from the average nodes population. We specialize our theory to the prototypical case of a linearly separable data and a linear hinge loss, for which the dynamics can be explicitly solved in the infinite dataset limit. This allows us to address in a simple setting several phenomena appearing in modern networks such as slowing down of training dynamics, crossover between rich and lazy learning, and overfitting. Finally, we assess the limitations of mean-field theory by studying the case of large but finite number of nodes and of training samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 被引用 94 次
- Towards Understanding the Condensation of Neural Networks at Initial TrainingHanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang 等NeurIPS 2022 · 被引用 42 次
- Empirical Phase Diagram for Three-layer Neural Networks with Infinite WidthHanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo 等NeurIPS 2022 · 被引用 22 次
- On the Equivalence between Neural Network and Support Vector MachineYilan Chen, Wei Huang, Lam M. Nguyen, Tsui-Wei WengNeurIPS 2021 · 被引用 21 次
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable DataNikita Tsoy, Nikola KonstantinovICML 2024 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer NetworksAndrea Montanari, Pierfrancesco UrbaniNeurIPS 2025 · 被引用 29 次
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 被引用 95 次
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala 等NeurIPS 2022 · 被引用 59 次
- Training shallow ReLU networks on noisy data using hinge loss: when do we overfit and is it benign?Erin George, Michael Murray, William Swartworth, Deanna NeedellNeurIPS 2023 · 被引用 9 次
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient DescentShota Imai, Sota Nishiyama, Masaaki ImaizumiICML 2026
