The Power and Limitation of Pretraining-Finetuning for Linear Regression under Covariate Shift
Jingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu, Sham M. Kakade
Abstract
We study linear regression under covariate shift, where the marginal distribution over the input covariates differs in the source and the target domains, while the conditional distribution of the output given the input covariates is similar across the two domains. We investigate a transfer learning approach with pretraining on the source data and finetuning based on the target data (both conducted by online SGD) for this problem. We establish sharp instance-dependent excess risk upper and lower bounds for this approach. Our bounds suggest that for a large class of linear regression instances, transfer learning with source data (and scarce or no target data) is as effective as supervised learning with target data. In addition, we show that finetuning, even with only a small amount of target data, could drastically reduce the amount of source data required by pretraining. Our theory sheds light on the effectiveness and limitation of pretraining as well as the benefits of finetuning for tackling covariate shift problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 836a1217-1fb8-4bfe-9c08-966e4032b3a6Cited by top-tier papers13
- Scaling Laws in Linear Regression: Compute, Parameters, and DataLicong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett et al.NeurIPS 2024 · 57 citations
- A Statistical Theory of Regularization-Based Continual LearningXuyang Zhao, Huiyuan Wang, Weiran Huang, Wei LinICML 2024 · 40 citations
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
- Demystifying Disagreement-on-the-Line in High DimensionsDonghwan Lee, Behrad Moniri, Xinmeng Huang, Edgar Dobriban et al.ICML 2023 · 12 citations
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear RegressionTingkai Yan, Haodong Wen, Binghui Li, Kairong Luo et al.ICLR 2026 · 12 citations
Builds on5
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 52 citations
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2021 · 41 citations
- A new similarity measure for covariate shift with applications to nonparametric regressionReese Pathak, Cong Ma, Martin J. WainwrightICML 2022 · 40 citations
- Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear RegressionJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu et al.ICML 2022 · 38 citations
Related papers
- Universality in Transfer Learning for Linear ModelsReza Ghane, Danil Akhtiamov, Babak HassibiNeurIPS 2024 · 8 citations
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 13 citations
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural NetworksSeyed Mohammadreza Mousavi Kalan, Zalan Fabian, Salman Avestimehr, Mahdi SoltanolkotabiNeurIPS 2020 · 37 citations
- Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index ModelsTaj Jones-McCormick, Aukosh Jagannath, Subhabrata SenICML 2025
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian et al.NeurIPS 2022 · 69 citations
