Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
Dan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry, Yu-Xiang Wang
摘要
We study the generalization of two-layer ReLU neural networks in a univariate nonparametric regression problem with noisy labels. This is a problem where kernels (e.g. NTK) are provably sub-optimal and benign overfitting does not happen, thus disqualifying existing theory for interpolating (0-loss, global optimal) solutions. We present a new theory of generalization for local minima that gradient descent with a constant learning rate can stably converge to. We show that gradient descent with a fixed learning rate can only find local minima that represent smooth functions with a certain weighted first order total variation bounded by where is the label noise level, is short for mean squared error against the ground truth, and hides a logarithmic factor. Under mild assumptions, we also prove a nearly-optimal MSE bound of within the strict interior of the support of the data points. Our theoretical results are validated by extensive simulation that demonstrates large learning rate training induces sparse linear spline fits. To the best of our knowledge, we are the first to obtain generalization bound via minima stability in the non-interpolation case and the first to show ReLU NNs without regularization can achieve near-optimal rates in nonparametric regression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Large Stepsizes Accelerate Gradient Descent for Regularized Logistic RegressionJingfeng Wu, Pierre Marion, Peter L. BartlettNeurIPS 2025 · 被引用 12 次
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 被引用 8 次
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 · 被引用 7 次
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva 等NeurIPS 2024 · 被引用 6 次
- Simplicity Bias and Optimization Threshold in Two-Layer ReLU NetworksEtienne Boursier, Nicolas FlammarionICML 2025
它引用的顶会 Paper25
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 被引用 172 次
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 被引用 140 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- On Linear Stability of SGD and Input-Smoothness of Neural NetworksChao Ma, Lexing YingNeurIPS 2021 · 被引用 73 次
- The Implicit Bias of Minima Stability: A View from Function SpaceRotem Mulayoff, Tomer Michaeli, Daniel SoudryNeurIPS 2021 · 被引用 65 次
相关 Paper
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 被引用 15 次
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello 等ICML 2023 · 被引用 12 次
- Optimal Rates for Generalization of Gradient Descent for Deep ReLU ClassificationYuanfan Li, Yunwen Lei, Zheng-Chu Guo, Yiming YingNeurIPS 2025 · 被引用 4 次
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 被引用 1 次
- Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?Kaiqi Zhang, Yu-Xiang WangICLR 2023 · 被引用 3 次
