Universality in Transfer Learning for Linear Models
Reza Ghane, Danil Akhtiamov, Babak Hassibi
Abstract
We study the problem of transfer learning and fine-tuning in linear models for both regression and binary classification. In particular, we consider the use of stochastic gradient descent (SGD) on a linear model initialized with pretrained weights and using a small training data set from the target distribution. In the asymptotic regime of large models, we provide an exact and rigorous analysis and relate the generalization errors (in regression) and classification errors (in binary classification) for the pretrained and fine-tuned models. In particular, we give conditions under which the fine-tuned model outperforms the pretrained one. An important aspect of our work is that all the results are"universal", in the sense that they depend only on the first and second order statistics of the target distribution. They thus extend well beyond the standard Gaussian assumptions commonly made in the literature. Furthermore, our universality results extend beyond standard SGD training to the test error of a classification task trained using a ridge regression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b99a06b1-31f6-4fcf-9d4d-4c4d1168b43aBuilds on4
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 78 citations
- Universality laws for Gaussian mixtures in generalized linear modelsYatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro et al.NeurIPS 2023 · 40 citations
- Deterministic equivalent and error universality of deep random features learningDominik Schröder, Hugo Cui, Daniil Dmitriev, Bruno LoureiroICML 2023 · 37 citations
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
Related papers
- The Power and Limitation of Pretraining-Finetuning for Linear Regression under Covariate ShiftJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2022 · 29 citations
- A Theoretical Analysis of the Test Error of Finite-Rank Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, Ivan Dokmanic et al.NeurIPS 2023 · 11 citations
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2021 · 41 citations
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
- Risk Bounds of Accelerated SGD for Overparameterized Linear RegressionXuheng Li, Yihe Deng, Jingfeng Wu, Dongruo Zhou et al.ICLR 2024 · 7 citations
