Generalization in Kernel Regression Under Realistic Assumptions
Daniel Barzilai, Ohad Shamir
摘要
It is by now well-established that modern over-parameterized models seem to elude the bias-variance tradeoff and generalize well despite overfitting noise. Many recent works attempt to analyze this phenomenon in the relatively tractable setting of kernel regression. However, as we argue in detail, most past works on this topic either make unrealistic assumptions, or focus on a narrow problem setup. This work aims to provide a unified theory to upper bound the excess risk of kernel regression for nearly all common and realistic settings. Specifically, we provide rigorous bounds that hold for common kernels and for any amount of regularization, noise, any input dimension, and any number of samples. Furthermore, we provide relative perturbation bounds for the eigenvalues of kernel matrices, which may be of independent interest. These reveal a self-regularization phenomenon, whereby a heavy tail in the eigendecomposition of the kernel provides it with an implicit form of regularization, enabling good generalization. When applied to common kernels, our results imply benign overfitting in high input dimensions, nearly tempered overfitting in fixed dimensions, and explicit convergence rates for regularized regression. As a by-product, we obtain time-dependent bounds for neural networks trained in the kernel regime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 被引用 13 次
- Characterizing Overfitting in Kernel Ridgeless Regression Through the EigenspectrumTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusICML 2024 · 被引用 12 次
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 被引用 9 次
- Provable Tempered Overfitting of Minimal Nets and Typical NetsItamar Harel, William Hoza, Gal Vardi, Itay Evron 等NeurIPS 2024 · 被引用 7 次
- A Comprehensive Analysis on the Learning Curve in Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusNeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper12
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 被引用 109 次
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 被引用 101 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
相关 Paper
- On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law DecayYicheng Li, Haobo Zhang, Qian LinNeurIPS 2023 · 被引用 23 次
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit 等NeurIPS 2022 · 被引用 53 次
- An Agnostic View on the Cost of Overfitting in (Kernel) Ridge RegressionLijia Zhou, James B. Simon, Gal Vardi, Nathan SrebroICLR 2024 · 被引用 2 次
- Strong inductive biases provably prevent harmless interpolationMichael Aerni, Marco Milanta, Konstantin Donhauser, Fanny YangICLR 2023
- Noisy Interpolation Learning with Shallow Univariate ReLU NetworksNirmit Joshi, Gal Vardi, Nathan SrebroICLR 2024 · 被引用 12 次
