Kernel Alignment Risk Estimator: Risk Prediction from Training Data
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel
Abstract
We study the risk (i.e. generalization error) of Kernel Ridge Regression (KRR) for a kernel with ridge and i.i.d. observations. For this, we introduce two objects: the Signal Capture Threshold (SCT) and the Kernel Alignment Risk Estimator (KARE). The SCT is a function of the data distribution: it can be used to identify the components of the data that the KRR predictor captures, and to approximate the (expected) KRR risk. This then leads to a KRR risk approximation by the KARE , an explicit function of the training data, agnostic of the true data distribution. We phrase the regression problem in a functional setting. The key results then follow from a finite-size analysis of the Stieltjes transform of general Wishart random matrices. Under a natural universality assumption (that the KRR moments depend asymptotically on the first two moments of the observations) we capture the mean and variance of the KRR predictor. We numerically investigate our findings on the Higgs and MNIST datasets for various classical kernels: the KARE gives an excellent approximation of the risk, thus supporting our universality assumption. Using the KARE, one can compare choices of Kernels and hyperparameters directly from the training set. The KARE thus provides a promising data-dependent procedure to select Kernels that generalize well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ed4076f-0b8b-4c97-b03c-193b47862df0Cited by top-tier papers31
- The Inductive Bias of Quantum KernelsJonas M. Kübler, Simon Buchholz, Bernhard SchölkopfNeurIPS 2021 · 190 citations
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 109 citations
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 90 citations
Builds on4
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Neural Kernels Without TangentsVaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil et al.ICML 2020 · 93 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 52 citations
Related papers
- Target alignment in truncated kernel ridge regressionArash A. Amini, Richard Baumgartner, Dai FengNeurIPS 2022 · 4 citations
- An Agnostic View on the Cost of Overfitting in (Kernel) Ridge RegressionLijia Zhou, James B. Simon, Gal Vardi, Nathan SrebroICLR 2024 · 2 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- A Comprehensive Analysis on the Learning Curve in Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusNeurIPS 2024 · 7 citations
- Precise Learning Curves and Higher-Order Scalings for Dot-product Kernel RegressionLechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu et al.NeurIPS 2022 · 27 citations
