Overfitting Can Be Harmless for Basis Pursuit, But Only to a Degree
Peizhong Ju, Xiaojun Lin, Jia Liu
摘要
Recently, there have been significant interests in studying the so-called "double-descent" of the generalization error of linear regression models under the overparameterized and overfitting regime, with the hope that such analysis may provide the first step towards understanding why overparameterized deep neural networks (DNN) still generalize well. However, to date most of these studies focused on the min -norm solution that overfits the data. In contrast, in this paper we study the overfitting solution that minimizes the -norm, which is known as Basis Pursuit (BP) in the compressed sensing literature. Under a sparse true linear regression model with i.i.d. Gaussian features, we show that for a large range of up to a limit that grows exponentially with the number of samples , with high probability the model error of BP is upper bounded by a value that decreases with . To the best of our knowledge, this is the first analytical result in the literature establishing the double-descent of overfitting BP for finite and . Further, our results reveal significant differences between the double-descent of BP and min -norm solutions. Specifically, the double-descent upper-bound of BP is independent of the signal strength, and for high SNR and sparse models the descent-floor of BP can be much lower and wider than that of min -norm solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 被引用 74 次
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 被引用 58 次
- Fast rates for noisy interpolation require rethinking the effect of inductive biasKonstantin Donhauser, Nicolò Ruggeri, Stefan Stojanovic, Fanny YangICML 2022 · 被引用 24 次
- On the Generalization Power of Overfitted Two-Layer Neural Tangent Kernel ModelsPeizhong Ju, Xiaojun Lin, Ness B. ShroffICML 2021 · 被引用 13 次
- Noisy Interpolation Learning with Shallow Univariate ReLU NetworksNirmit Joshi, Gal Vardi, Nathan SrebroICLR 2024 · 被引用 12 次
相关 Paper
- Theoretical Characterization of the Generalization Performance of Overfitted Meta-LearningPeizhong Ju, Yingbin Liang, Ness B. ShroffICLR 2023 · 被引用 3 次
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 被引用 81 次
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki 等NeurIPS 2021 · 被引用 12 次
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 被引用 94 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
