A Theoretical Analysis of Fine-tuning with Linear Teachers
Gal Shachaf, Alon Brutzkus, Amir Globerson
摘要
Fine-tuning is a common practice in deep learning, achieving excellent generalization results on downstream tasks using relatively little training data. Although widely used in practice, it is lacking strong theoretical understanding. We analyze the sample complexity of this scheme for regression with linear teachers in several architectures. Intuitively, the success of fine-tuning depends on the similarity between the source tasks and the target task, however measuring it is non trivial. We show that a relevant measure considers the relation between the source task, the target task and the covariance structure of the target data. In the setting of linear regression, we show that under realistic settings a substantial sample complexity reduction is plausible when the above measure is low. For deep linear regression, we present a novel result regarding the inductive bias of gradient-based training when the network is initialized with pretrained weights. Using this result we show that the similarity measure for this setting is also affected by the depth of the network. We further present results on shallow ReLU models, and analyze the dependence of sample complexity there on source and target tasks. We empirically demonstrate our results for both synthetic and realistic data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Exact learning dynamics of deep linear networks with prior knowledgeLukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. SaxeNeurIPS 2022 · 被引用 75 次
- Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization GuaranteesHaotian Ju, Dongyue Li, Hongyang R. ZhangICML 2022 · 被引用 41 次
- TSDS: Data Selection for Task-Specific Model FinetuningZifan Liu, Amin Karbasi, Theodoros RekatsinasNeurIPS 2024 · 被引用 35 次
- Implicit Regularization in Hierarchical Tensor Factorization and Deep Convolutional Neural NetworksNoam Razin, Asaf Maman, Nadav CohenICML 2022 · 被引用 34 次
- Active Multi-Task Representation LearningYifang Chen, Kevin Jamieson, Simon S. DuICML 2022 · 被引用 18 次
它引用的顶会 Paper14
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
相关 Paper
- Towards Sample-efficient Overparameterized Meta-learningYue Sun, Adhyyan Narang, Halil Ibrahim Gulluk, Samet Oymak 等NeurIPS 2021 · 被引用 26 次
- Features are fate: a theory of transfer learning in high-dimensional regressionJavan Tahir, Surya Ganguli, Grant M. RotskoffICML 2025
- Transfer Learning in Infinite Width Feature Learning NetworksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICLR 2026 · 被引用 2 次
- A Theory of How Pretraining Shapes Inductive Bias in Fine-TuningNicolas Anguita, Francesco Locatello, Andrew Saxe, Marco Mondelli 等ICML 2026 · 被引用 3 次
- Student Specialization in Deep Rectified Networks With Finite Width and Input DimensionYuandong TianICML 2020 · 被引用 12 次
