On the Universality of the Double Descent Peak in Ridgeless Regression
David Holzmüller
Abstract
We prove a non-asymptotic distribution-independent lower bound for the expected mean squared generalization error caused by label noise in ridgeless linear regression. Our lower bound generalizes a similar known result to the overparameterized (interpolating) regime. In contrast to most previous works, our analysis applies to a broad class of input distributions with almost surely full-rank feature matrices, which allows us to cover various types of deterministic or random feature maps. Our lower bound is asymptotically sharp and implies that in the presence of label noise, ridgeless linear regression does not perform well around the interpolation threshold for any of these feature maps. We analyze the imposed assumptions in detail and provide a theory for analytic (random) feature maps. Using this theory, we can show that our assumptions are satisfied for input distributions with a (Lebesgue) density and feature maps given by random deep neural networks with analytic activation functions like sigmoid, tanh, softplus, or GELU. As further examples, we show that feature maps from random Fourier features and polynomial kernels also satisfy our assumptions. We complement our theory with further experimental and analytic results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4933d85-6512-4c06-a7c1-37cb60cb7f1cCited by top-tier papers4
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit et al.NeurIPS 2022 · 53 citations
- Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimensionMoritz Haas, David Holzmüller, Ulrike von Luxburg, Ingo SteinwartNeurIPS 2023 · 30 citations
- Monotonicity and Double Descent in Uncertainty Estimation with Gaussian ProcessesLiam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. MahoneyICML 2023 · 9 citations
- Models of Heavy-Tailed Mechanistic UniversalityLiam Hodgkinson, Zhichao Wang, Michael W. MahoneyICML 2025
Builds on6
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 148 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 94 citations
Related papers
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methodsTaiji Suzuki, Shunta AkiyamaICLR 2021 · 12 citations
- Characterizing Overfitting in Kernel Ridgeless Regression Through the EigenspectrumTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusICML 2024 · 12 citations
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry et al.NeurIPS 2024 · 15 citations
- Asymptotic normality and confidence intervals for derivatives of 2-layers neural network in the random features modelYiwei Shen, Pierre C. BellecNeurIPS 2020 · 1 citation
