The Power and Limitation of Pretraining-Finetuning for Linear Regression under Covariate Shift
Jingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu, Sham M. Kakade
摘要
We study linear regression under covariate shift, where the marginal distribution over the input covariates differs in the source and the target domains, while the conditional distribution of the output given the input covariates is similar across the two domains. We investigate a transfer learning approach with pretraining on the source data and finetuning based on the target data (both conducted by online SGD) for this problem. We establish sharp instance-dependent excess risk upper and lower bounds for this approach. Our bounds suggest that for a large class of linear regression instances, transfer learning with source data (and scarce or no target data) is as effective as supervised learning with target data. In addition, we show that finetuning, even with only a small amount of target data, could drastically reduce the amount of source data required by pretraining. Our theory sheds light on the effectiveness and limitation of pretraining as well as the benefits of finetuning for tackling covariate shift problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Scaling Laws in Linear Regression: Compute, Parameters, and DataLicong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett 等NeurIPS 2024 · 被引用 57 次
- A Statistical Theory of Regularization-Based Continual LearningXuyang Zhao, Huiyuan Wang, Weiran Huang, Wei LinICML 2024 · 被引用 40 次
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 被引用 23 次
- Demystifying Disagreement-on-the-Line in High DimensionsDonghwan Lee, Behrad Moniri, Xinmeng Huang, Edgar Dobriban 等ICML 2023 · 被引用 12 次
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear RegressionTingkai Yan, Haodong Wen, Binghui Li, Kairong Luo 等ICLR 2026 · 被引用 12 次
它引用的顶会 Paper5
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 被引用 45 次
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2021 · 被引用 41 次
- A new similarity measure for covariate shift with applications to nonparametric regressionReese Pathak, Cong Ma, Martin J. WainwrightICML 2022 · 被引用 40 次
- Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear RegressionJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu 等ICML 2022 · 被引用 38 次
相关 Paper
- Universality in Transfer Learning for Linear ModelsReza Ghane, Danil Akhtiamov, Babak HassibiNeurIPS 2024 · 被引用 8 次
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 被引用 13 次
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural NetworksSeyed Mohammadreza Mousavi Kalan, Zalan Fabian, Salman Avestimehr, Mahdi SoltanolkotabiNeurIPS 2020 · 被引用 37 次
- Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index ModelsTaj Jones-McCormick, Aukosh Jagannath, Subhabrata SenICML 2025
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian 等NeurIPS 2022 · 被引用 69 次
