Features are fate: a theory of transfer learning in high-dimensional regression
Javan Tahir, Surya Ganguli, Grant M. Rotskoff
Abstract
With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of "task similarity" is still lacking. We adopt a feature-centric viewpoint on transfer learning and establish a number of theoretical results that demonstrate that when the target task is well represented by the feature space of the pre-trained model, transfer learning outperforms training from scratch. We study deep linear networks as a minimal model of transfer learning in which we can analytically characterize the transferability phase diagram as a function of the target dataset size and the feature space overlap. For this model, we establish rigorously that when the feature space overlap between the source and target tasks is sufficiently strong, both linear transfer and fine-tuning improve performance, especially in the low data limit. These results build on an emerging understanding of feature learning dynamics in deep linear networks, and we demonstrate numerically that the rigorous results we derive for the linear case also apply to nonlinear networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Multifidelity Simulation-based Inference for Computationally Expensive SimulatorsAnastasia Nastya Krouglova, Hayden R. Johnson, Basile Confavreux, Michael Deistler et al.ICLR 2026 · 17 citations
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada et al.NeurIPS 2025 · 15 citations
- A Theory of How Pretraining Shapes Inductive Bias in Fine-TuningNicolas Anguita, Francesco Locatello, Andrew Saxe, Marco Mondelli et al.ICML 2026 · 3 citations
- Transfer Learning for Benign Overfitting in High-Dimensional Linear RegressionYeichan Kim, Ilmun Kim, Seyoung ParkNeurIPS 2025 · 2 citations
- Transfer Learning in Infinite Width Feature Learning NetworksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICLR 2026 · 2 citations
Builds on14
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 267 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
Related papers
- A Theoretical Analysis of Fine-tuning with Linear TeachersGal Shachaf, Alon Brutzkus, Amir GlobersonNeurIPS 2021 · 23 citations
- Adversarially-Trained Deep Nets Transfer Better: Illustration on Image ClassificationFrancisco Utrera, Evan Kravitz, N. Benjamin Erichson, Rajiv Khanna et al.ICLR 2021 · 42 citations
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural NetworksSeyed Mohammadreza Mousavi Kalan, Zalan Fabian, Salman Avestimehr, Mahdi SoltanolkotabiNeurIPS 2020 · 37 citations
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian et al.NeurIPS 2022 · 69 citations
