Triple descent and the two kinds of overfitting: where & why do they appear?
Stéphane d'Ascoli, Levent Sagun, Giulio Biroli
Abstract
A recent line of research has highlighted the existence of a "double descent" phenomenon in deep learning, whereby increasing the number of training examples N causes the generalization error of neural networks to peak when N is of the same order as the number of parameters P . In earlier works, a similar phenomenon was shown to exist in simpler models such as linear regression, where the peak instead occurs when N is equal to the input dimension D. Since both peaks coincide with the interpolation threshold, they are often conflated in the litterature. In this paper, we show that despite their apparent similarity, these two scenarios are inherently different. In fact, both peaks can co-exist when neural networks are applied to noisy regression tasks. The relative size of the peaks is then governed by the degree of nonlinearity of the activation function. Building on recent developments in the analysis of random feature models, we provide a theoretical ground for this sample-wise triple descent. As shown previously, the nonlinear peak at N = P is a true divergence caused by the extreme sensitivity of the output function to both the noise corrupting the labels and the initialization of the random features (or the weights in neural networks). This peak survives in the absence of noise, but can be suppressed by regularization. In contrast, the linear peak at N = D is solely due to overfitting the noise in the labels, and forms earlier during training. We show that this peak is implicitly regularized by the nonlinearity, which is why it only becomes salient at high noise and is weakly affected by explicit regularization. Throughout the paper, we compare analytical results obtained in the random feature model with the outcomes of numerical experiments involving deep neural networks. 1 Also called the jamming peak due to similarities with a well-studied phenomenon in the Statistical Physics literature [14, 15, 16, 17, 18] .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f80dac2b-740f-4784-8cee-e17b695353b4Cited by top-tier papers9
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit et al.NeurIPS 2022 · 53 citations
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
- On the Universality of the Double Descent Peak in Ridgeless RegressionDavid HolzmüllerICLR 2021 · 16 citations
- Learning Curves for Deep Structured Gaussian Feature ModelsJacob A. Zavatone-Veth, Cengiz PehlevanNeurIPS 2023 · 15 citations
- Information bottleneck theory of high-dimensional regression: relevancy, efficiency and optimalityVudtiwat Ngampruetikorn, David J. SchwabNeurIPS 2022 · 11 citations
Builds on9
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 148 citations
Related papers
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki et al.NeurIPS 2021 · 12 citations
- Overfitting Can Be Harmless for Basis Pursuit, But Only to a DegreePeizhong Ju, Xiaojun Lin, Jia LiuNeurIPS 2020 · 20 citations
- Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature modelAntoine Bodin, Nicolas MacrisNeurIPS 2021 · 19 citations
- Asymptotic normality and confidence intervals for derivatives of 2-layers neural network in the random features modelYiwei Shen, Pierre C. BellecNeurIPS 2020 · 1 citation
