Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension
Moritz Haas, David Holzmüller, Ulrike von Luxburg, Ingo Steinwart
Abstract
The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has been established for some learning methods, current literature suggests that for regression with typical kernel methods and wide neural networks, benign overfitting requires a high-dimensional setting where the dimension grows with the sample size. In this paper, we show that the smoothness of the estimators, and not the dimension, is the key: benign overfitting is possible if and only if the estimator's derivatives are large enough. We generalize existing inconsistency results to non-interpolating models and more kernels to show that benign overfitting with moderate derivatives is impossible in fixed dimension. Conversely, we show that rate-optimal benign overfitting is possible for regression with a sequence of spiky-smooth kernels with large derivatives. Using neural tangent kernels, we translate our results to wide neural networks. We prove that while infinite-width networks do not overfit benignly with the ReLU activation, this can be fixed by adding small high-frequency fluctuations to the activation function. Our experiments verify that such neural networks, while overfitting, can indeed generalize well even on low-dimensional data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16cea452-3bc2-48d3-aa82-516ca7251834Cited by top-tier papers13
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry et al.NeurIPS 2024 · 15 citations
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 13 citations
- Characterizing Overfitting in Kernel Ridgeless Regression Through the EigenspectrumTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusICML 2024 · 12 citations
- Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondAlan Jeffares, Alicia Curth, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 9 citations
Builds on14
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Early-stopped neural networks are consistentZiwei Ji, Justin D. Li, Matus TelgarskyNeurIPS 2021 · 58 citations
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 41 citations
- Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and InitializationMariia Seleznova, Gitta KutyniokICML 2022 · 34 citations
Related papers
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit et al.NeurIPS 2022 · 53 citations
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- Benign Overfitting in Two-layer ReLU Convolutional Neural NetworksYiwen Kou, Zixiang Chen, Yuanzhou Chen, Quanquan GuICML 2023 · 32 citations
- Strong inductive biases provably prevent harmless interpolationMichael Aerni, Marco Milanta, Konstantin Donhauser, Fanny YangICLR 2023
- From Tempered to Benign Overfitting in ReLU Neural NetworksGuy Kornowski, Gilad Yehudai, Ohad ShamirNeurIPS 2023 · 18 citations
