A Sample Complexity Separation between Non-Convex and Convex Meta-Learning
Nikunj Saunshi, Yi Zhang, Mikhail Khodak, Sanjeev Arora
Abstract
One popular trend in meta-learning is to learn from many training tasks a common initialization for a gradient-based method that can be used to solve a new task with few samples. The theory of meta-learning is still in its early stages, with several recent learning-theoretic analyses of methods such as Reptile [Nichol et al., 2018] being for convex models. This work shows that convex-case analysis might be insufficient to understand the success of meta-learning, and that even for non-convex models it is important to look inside the optimization black-box, specifically at properties of the optimization trajectory. We construct a simple meta-learning instance that captures the problem of one-dimensional subspace learning. For the convex formulation of linear regression on this instance, we show that the new task sample complexity of any initialization-based meta-learning algorithm is , where is the input dimension. In contrast, for the non-convex formulation of a two layer linear network on the same instance, we show that both Reptile and multi-task representation learning can have new task sample complexity of , demonstrating a separation from convex meta-learning. Crucially, analyses of the training dynamics of these methods reveal that they can meta-learn the correct subspace onto which the data should be projected.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective AdaptationHaoxiang Wang, Han Zhao, Bo LiICML 2021 · 108 citations
- How Important is the Train-Validation Split in Meta-Learning?Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao et al.ICML 2021 · 60 citations
- How Fine-Tuning Allows for Effective Meta-LearningKurtland Chua, Qi Lei, Jason D. LeeNeurIPS 2021 · 57 citations
- MAML and ANIL Provably Learn RepresentationsLiam Collins, Aryan Mokhtari, Sewoong Oh, Sanjay ShakkottaiICML 2022 · 38 citations
- Subspace Learning for Effective Meta-LearningWeisen Jiang, James T. Kwok, Yu ZhangICML 2022 · 28 citations
Related papers
- A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-LearningNikunj Saunshi, Arushi Gupta, Wei HuICML 2021 · 19 citations
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 218 citations
- Statistically and Computationally Efficient Linear Meta-representation LearningKiran Koshy Thekumparampil, Prateek Jain, Praneeth Netrapalli, Sewoong OhNeurIPS 2021 · 28 citations
- Task Relatedness-Based Generalization Bounds for Meta LearningJiechao Guan, Zhiwu LuICLR 2022 · 11 citations
- The Effect of Diversity in Meta-LearningRamnath Kumar, Tristan Deleu, Yoshua BengioAAAI 2023 · 18 citations
