From Tempered to Benign Overfitting in ReLU Neural Networks
Guy Kornowski, Gilad Yehudai, Ohad Shamir
摘要
Overparameterized neural networks (NNs) are observed to generalize well even when trained to perfectly fit noisy data. This phenomenon motivated a large body of work on "benign overfitting", where interpolating predictors achieve near-optimal performance. Recently, it was conjectured and empirically observed that the behavior of NNs is often better described as "tempered overfitting", where the performance is non-optimal yet also non-trivial, and degrades as a function of the noise level. However, a theoretical justification of this claim for non-linear NNs has been lacking so far. In this work, we provide several results that aim at bridging these complementing views. We study a simple classification setting with 2-layer ReLU NNs, and prove that under various assumptions, the type of overfitting transitions from tempered in the extreme case of one-dimensional data, to benign in high dimensions. Thus, we show that the input dimension has a crucial role on the overfitting profile in this setting, which we also validate empirically for intermediate dimensions. Overall, our results shed light on the intricate connections between the dimension, sample size, architecture and training algorithm on the one hand, and the type of resulting overfitting on the other hand. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Benign Overfitting and Grokking in ReLU Networks for XOR Cluster DataZhiwei Xu, Yutong Wang, Spencer Frei, Gal Vardi 等ICLR 2024 · 被引用 39 次
- The Double-Edged Sword of Implicit Bias: Generalization vs. Robustness in ReLU NetworksSpencer Frei, Gal Vardi, Peter L. Bartlett, Nati SrebroNeurIPS 2023 · 被引用 25 次
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry 等NeurIPS 2024 · 被引用 15 次
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 被引用 15 次
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 被引用 13 次
它引用的顶会 Paper15
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Benign Overfitting in Two-layer Convolutional Neural NetworksYuan Cao, Zixiang Chen, Misha Belkin, Quanquan GuNeurIPS 2022 · 被引用 121 次
- In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating PredictorsJeffrey Negrea, Gintare Karolina Dziugaite, Daniel M. RoyICML 2020 · 被引用 66 次
- Uniform Convergence of Interpolators: Gaussian Width, Norm Bounds and Benign OverfittingFrederic Koehler, Lijia Zhou, Danica J. Sutherland, Nathan SrebroNeurIPS 2021 · 被引用 65 次
相关 Paper
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit 等NeurIPS 2022 · 被引用 53 次
- Noisy Interpolation Learning with Shallow Univariate ReLU NetworksNirmit Joshi, Gal Vardi, Nathan SrebroICLR 2024 · 被引用 12 次
- Provable Tempered Overfitting of Minimal Nets and Typical NetsItamar Harel, William Hoza, Gal Vardi, Itay Evron 等NeurIPS 2024 · 被引用 7 次
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 被引用 3 次
- Benign overfitting in leaky ReLU networks with moderate input dimensionKedar Karhadkar, Erin George, Michael Murray, Guido F. Montúfar 等NeurIPS 2024 · 被引用 5 次
