An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini, Giulio Biroli
Abstract
Neural networks have been shown to perform incredibly well in classification tasks over structured high-dimensional datasets. However, the learning dynamics of such networks is still poorly understood. In this paper we study in detail the training dynamics of a simple type of neural network: a single hidden layer trained to perform a classification task. We show that in a suitable mean-field limit this case maps to a single-node learning problem with a time-dependent dataset determined self-consistently from the average nodes population. We specialize our theory to the prototypical case of a linearly separable data and a linear hinge loss, for which the dynamics can be explicitly solved in the infinite dataset limit. This allows us to address in a simple setting several phenomena appearing in modern networks such as slowing down of training dynamics, crossover between rich and lazy learning, and overfitting. Finally, we assess the limitations of mean-field theory by studying the case of large but finite number of nodes and of training samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb045d64-d85c-4e08-b04d-c8d9cec795bcCited by top-tier papers6
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
- Towards Understanding the Condensation of Neural Networks at Initial TrainingHanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang et al.NeurIPS 2022 · 42 citations
- Empirical Phase Diagram for Three-layer Neural Networks with Infinite WidthHanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo et al.NeurIPS 2022 · 22 citations
- On the Equivalence between Neural Network and Support Vector MachineYilan Chen, Wei Huang, Lam M. Nguyen, Tsui-Wei WengNeurIPS 2021 · 21 citations
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable DataNikita Tsoy, Nikola KonstantinovICML 2024 · 12 citations
Builds on1
Related papers
- Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer NetworksAndrea Montanari, Pierfrancesco UrbaniNeurIPS 2025 · 29 citations
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 95 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- Training shallow ReLU networks on noisy data using hinge loss: when do we overfit and is it benign?Erin George, Michael Murray, William Swartworth, Deanna NeedellNeurIPS 2023 · 9 citations
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient DescentShota Imai, Sota Nishiyama, Masaaki ImaizumiICML 2026
