Computational Efficiency under Covariate Shift in Kernel Ridge Regression
Andrea Della Vecchia, Arnaud Mavakala Watusadisi, Ernesto De Vito, Lorenzo Rosasco
Abstract
This paper addresses the covariate shift problem in the context of nonparametric regression within reproducing kernel Hilbert spaces (RKHSs). Covariate shift arises in supervised learning when the input distributions of the training and test data differ, presenting additional challenges for learning. Although kernel methods have optimal statistical properties, their high computational demands in terms of time and, particularly, memory, limit their scalability to large datasets. To address this limitation, the main focus of this paper is to explore the trade-off between computational efficiency and statistical accuracy under covariate shift. We investigate the use of random projections where the hypothesis space consists of a random subspace within a given RKHS. Our results show that, even in the presence of covariate shift, significant computational savings can be achieved without compromising learning performance. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
importance weighting (IW) function. This function corresponds to the Radon-Nikodym derivative of the test marginal distribution with respect to the training marginal distribution Huang et al. (2006); Sugiyama et al. (2012); Fang et al. (2020). Shimodaira (2000) was the first to demonstrate the consistency of the importance-weighted maximum likelihood estimator, while in Cortes et al. (2010) the authors derived suboptimal finite-sample bounds for restricted function classes with finite pseudodimension. In the context of nonparametric regression in reproducing kernel Hilbert spaces (RKHSs) and under the assumption that the regression function belongs to the RKHS (well-specified case), Gizewski et al. (2022) provided optimal excess risk convergence results for the minimizer of the reweighted empirical risk when the weight function is known and uniformly bounded. In Ma et al. 2023, authors recently showed that the standard unweighted kernel ridge regression estimator is still minimax optimal under an appropriate choice of the regularization parameter, provided the IW function is uniformly bounded or its second moment is bounded. Similar results for nonparametric classification were obtained in Kpotufe & Martinet (2021). We also mention Schmidt-Hieber & Zamolodtchikov ( 2024); Pathak et al. (2022) in the context of nonparametric regression for classes of functions different from RKHSs, and Wen et al. (2014); Lei et al. (2021); Yamazaki et al. (2007) for parametric models. Building on Ma et al. (2023), in Gogolashvili et al. ( 2023) the authors extended the analysis to (simplified) misspecified case, where the regression function is not assumed to lie in the RKHS itself, but its projection is. In this setting, the authors showed that, under covariate shift, the unweighted classic KRR predictor is not a consistent estimator of the projection of the regression function. In such cases, IW correction is necessary.
Kernel methods provide a robust framework for nonparametric learning, but their scalability is limited by high computational and memory costs-challenges that clearly do not disappear under covariate shift. To address this, researchers have developed more efficient strategies, ranging from improved optimization (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 179 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- Domain Adaptation for Time Series Under Feature and Label ShiftsHuan He, Owen Queen, Teddy Koker, Consuelo Cuevas et al.ICML 2023 · 121 citations
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
Related papers
- Towards a Unified Analysis of Kernel-based Methods Under Covariate ShiftXingdong Feng, Xin He, Caixing Wang, Chao Wang et al.NeurIPS 2023 · 17 citations
- High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit RegularizationYihang Chen, Fanghui Liu, Taiji Suzuki, Volkan CevherICML 2024 · 5 citations
- ReTaSA: A Nonparametric Functional Estimation Approach for Addressing Continuous Target ShiftHwanwoo Kim, Xin Zhang, Jiwei Zhao, Qinglong TianICLR 2024 · 3 citations
- Robust Learning with the Hilbert-Schmidt Independence CriterionDaniel Greenfeld, Uri ShalitICML 2020 · 73 citations
- Double-Weighting for Covariate Shift AdaptationJosé Ignacio Segovia-Martín, Santiago Mazuelas, Anqi LiuICML 2023 · 9 citations
