Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization
Benjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka Zdeborová
摘要
We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer neural network with random iid inputs. We study the generalization performances of standard classifiers in the high-dimensional regime where is kept finite in the limit of a high dimension and number of samples . Our contribution is three-fold: First, we prove a formula for the generalization error achieved by regularized classifiers that minimize a convex loss. This formula was first obtained by the heuristic replica method of statistical physics. Secondly, focussing on commonly used loss functions and optimizing the regularization strength, we observe that while ridge regression performance is poor, logistic and hinge regression are surprisingly able to approach the Bayes-optimal generalization error extremely closely. As they lead to Bayes-optimal rates, a fact that does not follow from predictions of margin-based generalization error bounds. Third, we design an optimal loss and regularizer that provably leads to Bayes-optimal generalization error.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Label-Imbalanced and Group-Sensitive Classification under OverparameterizationGanesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, Christos ThrampoulidisNeurIPS 2021 · 被引用 122 次
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 被引用 94 次
- Benign Overfitting in Multiclass Classification: All Roads Lead to InterpolationKe Wang, Vidya Muthukumar, Christos ThrampoulidisNeurIPS 2021 · 被引用 56 次
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 被引用 49 次
它引用的顶会 Paper2
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
相关 Paper
- The Role of Regularization in Classification of High-dimensional Noisy Gaussian MixtureFrancesca Mignacco, Florent Krzakala, Yue M. Lu, Pierfrancesco Urbani 等ICML 2020 · 被引用 98 次
- Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regimeKeita Suzuki, Taiji SuzukiICLR 2024 · 被引用 2 次
- A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitionsGabriel Mel, Surya GanguliICML 2021 · 被引用 24 次
- Universal Consistency of Wide and Deep ReLU Neural Networks and Minimax Optimal Convergence Rates for Kolmogorov-Donoho Optimal Function ClassesHyunouk Ko, Xiaoming HuoICML 2024 · 被引用 1 次
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
