Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
Rustem Takhanov, Zhenisbek Assylbekov
Abstract
Conditionally positive definite (CPD) kernels are defined with respect to a function class . It is well known that such a kernel is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method --- called conditional kernel ridge regression (conditional KRR) due to its analogy with KRR --- where the estimated regression function is penalized by the square of its native space norm. This method is of interest because it can be viewed as classical linear regression, with features specified by , followed by the application of standard KRR to the residual (unexplained) component of the target variable. Methods of this type have recently attracted increasing attention. We study the statistical properties of this method by reducing its behavior to that of KRR with another fixed kernel, called the residual kernel. Our main theoretical result shows that such a reduction is indeed possible, at the cost of an additional term in the expected test risk, bounded by , where is the sample size and the hidden constant depends on the class and the input distribution. This reduction enables us to analyze conditional KRR in the case where is positive definite and is given by the first principal eigenfunctions in the Mercer decomposition of . We also consider the setting where consists of random features from a random feature representation of . It turns out that these two settings are closely related. Both our theoretical analysis and experiments confirm that conditional KRR outperforms standard KRR in these cases whenever the -component of the regression function is more pronounced than the residual part.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4da6d57-73d1-4c3d-b2db-358d7d8b9b89Builds on6
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 148 citations
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 109 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.NeurIPS 2020 · 74 citations
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit et al.NeurIPS 2022 · 53 citations
Related papers
- Quasi-Monte Carlo Features for Kernel ApproximationZhen Huang, Jiajin Sun, Yian HuangICML 2024 · 6 citations
- On the Saturation Effect of Kernel Ridge RegressionYicheng Li, Haobo Zhang, Qian LinICLR 2023 · 2 citations
- On the Size and Approximation Error of Distilled DatasetsAlaa Maalouf, Murad Tukan, Noel Loo, Ramin M. Hasani et al.NeurIPS 2023 · 2 citations
- A Comprehensive Analysis on the Learning Curve in Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusNeurIPS 2024 · 7 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
