Optimal Regularization can Mitigate Double Descent
Preetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu Ma
Abstract
Recent empirical and theoretical studies have shown that many learning algorithms -from linear regression to neural networks -can have test performance that is non-monotonic in quantities such the sample size and model size. This striking phenomenon, often referred to as "double descent", has raised questions of if we need to re-think our current understanding of generalization. In this work, we study whether the double-descent phenomenon can be avoided by using optimal regularization. Theoretically, we prove that for certain linear regression models with isotropic data distribution, optimally-tuned 2 regularization achieves monotonic test performance as we grow either the sample size or the model size. We also demonstrate empirically that optimally-tuned 2 regularization can mitigate double descent for more general models, including neural networks. Our results suggest that it may also be informative to study the test risk scalings of various algorithms in the context of appropriately tuned regularization. Recent works have demonstrated a ubiquitous "double descent" phenomenon present in a range of machine learning models, including decision trees, random features, linear regression, and deep neural networks (Opper
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f1cbaf5-6502-4808-b9c5-e07eefa3a3ccCited by top-tier papers39
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 94 citations
- Data Feedback Loops: Model-driven Amplification of Dataset BiasesRohan Taori, Tatsunori HashimotoICML 2023 · 67 citations
- Direct Parameterization of Lipschitz-Bounded Deep NetworksRuigang Wang, Ian R. ManchesterICML 2023 · 66 citations
Builds on2
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 81 citations
Related papers
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki et al.NeurIPS 2021 · 12 citations
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu et al.ICLR 2020 · 77 citations
- A U-turn on Double Descent: Rethinking Parameter Counting in Statistical LearningAlicia Curth, Alan Jeffares, Mihaela van der SchaarNeurIPS 2023 · 42 citations
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan et al.ICML 2020 · 40 citations
