On the Theory of Transfer Learning: The Importance of Task Diversity
Nilesh Tripuraneni, Michael I. Jordan, Chi Jin
Abstract
We provide new statistical guarantees for transfer learning via representation learning--when transfer is achieved by learning a feature representation shared across different tasks. This enables learning on new tasks using far less data than is required to learn them in isolation. Formally, we consider tasks parameterized by functions of the form in a general function class , where each is a task-specific function in and is the shared representation in . Letting denote the complexity measure of the function class, we show that for diverse training tasks (1) the sample complexity needed to learn the shared representation across the first training tasks scales as , despite no explicit access to a signal from the feature representation and (2) with an accurate estimate of the representation, the sample complexity needed to learn a new task scales only with . Our results depend upon a new general notion of task diversity--applicable to models with general tasks, features, and losses--as well as a novel chain rule for Gaussian complexities. Finally, we exhibit the utility of our general framework in several models of importance in the literature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f2ce528-1ec7-49ba-bed2-668587a22469Cited by top-tier papers100
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 1,081 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Transformers as Algorithms: Generalization and Stability in In-context LearningYingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, Samet OymakICML 2023 · 242 citations
Builds on1
Related papers
- Guarantees for Nonlinear Representation Learning: Non-identical Covariates, Dependent Data, Fewer SamplesThomas T. C. K. Zhang, Bruce D. Lee, Ingvar M. Ziemann, George J. Pappas et al.ICML 2024 · 2 citations
- Optimistic Rates for Multi-Task Representation LearningAustin Watkins, Enayat Ullah, Thanh Nguyen-Tang, Raman AroraNeurIPS 2023 · 12 citations
- Representation Learning Beyond Linear Prediction FunctionsZiping Xu, Ambuj TewariNeurIPS 2021 · 27 citations
- Evaluated CMI Bounds for Meta Learning: Tightness and ExpressivenessFredrik Hellström, Giuseppe DurisiNeurIPS 2022 · 15 citations
- Few-Shot Learning via Learning the Representation, ProvablySimon Shaolei Du, Wei Hu, Sham M. Kakade, Jason D. Lee et al.ICLR 2021 · 56 citations
