A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-Learning
Nikunj Saunshi, Arushi Gupta, Wei Hu
摘要
An effective approach in meta-learning is to utilize multiple "train tasks" to learn a good initialization for model parameters that can help solve unseen "test tasks" with very few samples by fine-tuning from this initialization. Although successful in practice, theoretical understanding of such methods is limited. This work studies an important aspect of these methods: splitting the data from each task into train (support) and validation (query) sets during meta-training. Inspired by recent work [Raghu et al., 2020] , we view such meta-learning methods through the lens of representation learning and argue that the train-validation split encourages the learned representation to be lowrank without compromising on expressivity, as opposed to the non-splitting variant that encourages high-rank representations. Since sample efficiency benefits from low-rankness, the splitting strategy will require very few samples to solve unseen test tasks. We present theoretical results that formalize this idea for linear representation learning on a subspace meta-learning instance, and experimentally verify this practical benefit of splitting in simulations and on standard meta-learning benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Generalization Bounds For Meta-Learning: An Information-Theoretic AnalysisQi Chen, Changjian Shui, Mario MarchandNeurIPS 2021 · 被引用 66 次
- MAML and ANIL Provably Learn RepresentationsLiam Collins, Aryan Mokhtari, Sewoong Oh, Sanjay ShakkottaiICML 2022 · 被引用 38 次
- ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text ClassificationHazel H. Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha 等AAAI 2022 · 被引用 33 次
- Understanding Benign Overfitting in Gradient-Based Meta LearningLisha Chen, Songtao Lu, Tianyi ChenNeurIPS 2022 · 被引用 20 次
它引用的顶会 Paper8
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 被引用 218 次
- Unraveling Meta-Learning: Understanding Feature Representations for Few-Shot TasksMicah Goldblum, Steven Reich, Liam Fowl, Renkun Ni 等ICML 2020 · 被引用 82 次
- Meta-learning for Mixed Linear RegressionWeihao Kong, Raghav Somani, Zhao Song, Sham M. Kakade 等ICML 2020 · 被引用 70 次
相关 Paper
- How Important is the Train-Validation Split in Meta-Learning?Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao 等ICML 2021 · 被引用 60 次
- Understanding Train-Validation Split in Meta-Learning with Neural NetworksXinzhe Zuo, Zixiang Chen, Huaxiu Yao, Yuan Cao 等ICLR 2023
- Provably Efficient Multi-Task Meta Bandit Learning via Shared RepresentationsJiabin Lin, Shana MoothedathNeurIPS 2025 · 被引用 2 次
- First-order ANIL provably learns representations despite overparametrisationOguz Kaan Yüksel, Etienne Boursier, Nicolas FlammarionICLR 2024 · 被引用 7 次
- A Sample Complexity Separation between Non-Convex and Convex Meta-LearningNikunj Saunshi, Yi Zhang, Mikhail Khodak, Sanjeev AroraICML 2020 · 被引用 32 次
