Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization
Benjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka Zdeborová
Abstract
We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer neural network with random iid inputs. We study the generalization performances of standard classifiers in the high-dimensional regime where is kept finite in the limit of a high dimension and number of samples . Our contribution is three-fold: First, we prove a formula for the generalization error achieved by regularized classifiers that minimize a convex loss. This formula was first obtained by the heuristic replica method of statistical physics. Secondly, focussing on commonly used loss functions and optimizing the regularization strength, we observe that while ridge regression performance is poor, logistic and hinge regression are surprisingly able to approach the Bayes-optimal generalization error extremely closely. As they lead to Bayes-optimal rates, a fact that does not follow from predictions of margin-based generalization error bounds. Third, we design an optimal loss and regularizer that provably leads to Bayes-optimal generalization error.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b689e9ac-6f42-4cb7-9d5b-7856da48d518Cited by top-tier papers15
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Label-Imbalanced and Group-Sensitive Classification under OverparameterizationGanesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, Christos ThrampoulidisNeurIPS 2021 · 122 citations
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 94 citations
- Benign Overfitting in Multiclass Classification: All Roads Lead to InterpolationKe Wang, Vidya Muthukumar, Christos ThrampoulidisNeurIPS 2021 · 56 citations
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 49 citations
Builds on2
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
Related papers
- The Role of Regularization in Classification of High-dimensional Noisy Gaussian MixtureFrancesca Mignacco, Florent Krzakala, Yue M. Lu, Pierfrancesco Urbani et al.ICML 2020 · 98 citations
- Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regimeKeita Suzuki, Taiji SuzukiICLR 2024 · 2 citations
- A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitionsGabriel Mel, Surya GanguliICML 2021 · 24 citations
- Universal Consistency of Wide and Deep ReLU Neural Networks and Minimax Optimal Convergence Rates for Kolmogorov-Donoho Optimal Function ClassesHyunouk Ko, Xiaoming HuoICML 2024 · 1 citation
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro et al.ICLR 2026 · 7 citations
