Transfer Learning with Affine Model Transformation
Shunya Minami, Kenji Fukumizu, Yoshihiro Hayashi, Ryo Yoshida
Abstract
Supervised transfer learning has received considerable attention due to its potential to boost the predictive power of machine learning in scenarios where data are scarce. Generally, a given set of source models and a dataset from a target domain are used to adapt the pre-trained models to a target domain by statistically learning domain shift and domain-specific factors. While such procedurally and intuitively plausible methods have achieved great success in a wide range of real-world applications, the lack of a theoretical basis hinders further methodological development. This paper presents a general class of transfer learning regression called affine model transfer, following the principle of expected-square loss minimization. It is shown that the affine model transfer broadly encompasses various existing methods, including the most common procedure based on neural feature extractors. Furthermore, the current paper clarifies theoretical properties of the affine model transfer such as generalization error and excess risk. Through several case studies, we demonstrate the practical benefits of modeling and estimating inter-domain commonality and domain-specific factors separately with the affine-type transfer models. *1.09 ± 0.232 *0.969 ± 0.144 *0.927 ± 0.170 Augmented 2.47 ± 0.406 1.90 ± 0.515 1.67 ± 0.552 *1.31 ± 0.214 1.16 ± 0.225 *0.984 ± 0.149 *0.897 ± 0.138 HTL-offset 2.29 ± 0.621 *1.69 ± 0.507 *1.49 ± 0.513 *1.22 ± 0.269 *1.09 ± 0.233 *0.969 ± 0.144 *0.925 ± 0.171 HTL-scale 2.32 ± 0.599 *1.71 ± 0.516 1.51 ± 0.513 *1.24 ± 0.271 *1.12 ± 0.234 *0.999 ± 0.175 0.948 ± 0.172 AffineTL-full *2.23 ± 0.554 *1.71 ± 0.501 *1.45 ± 0.458 *1.21 ± 0.256 *1.06 ± 0.219 *0.974 ± 0.164 *0.870 ± 0.121 AffineTL-const *2.30 ± 0.565 *1.73 ± 0.420 *1.48 ± 0.527 *1.20 ± 0.243 *1.04 ± 0.217 *0.963 ± 0.161 *0.884 ± 0.136
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
- Coupling-based Invertible Neural Networks Are Universal Diffeomorphism ApproximatorsTakeshi Teshima, Isao Ishikawa, Koichi Tojo, Kenta Oono et al.NeurIPS 2020 · 129 citations
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsAmanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey et al.EMNLP 2023 · 45 citations
- PAC-Net: A Model Pruning Approach to Inductive Transfer LearningSanghoon Myung, In Huh, Wonik Jang, Jae Myung Choe et al.ICML 2022 · 18 citations
Related papers
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural NetworksSeyed Mohammadreza Mousavi Kalan, Zalan Fabian, Salman Avestimehr, Mahdi SoltanolkotabiNeurIPS 2020 · 37 citations
- A General Class of Transfer Learning Regression without Implementation CostShunya Minami, Song Liu, Stephen Wu, Kenji Fukumizu et al.AAAI 2021 · 8 citations
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
- Adversarial Training Helps Transfer Learning via Better RepresentationsZhun Deng, Linjun Zhang, Kailas Vodrahalli, Kenji Kawaguchi et al.NeurIPS 2021 · 60 citations
- Test-time Adaptation for Regression by Subspace AlignmentKazuki Adachi, Shin'ya Yamaguchi, Atsutoshi Kumagai, Tomoki HamagamiICLR 2025
