How Important is the Train-Validation Split in Meta-Learning?
Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao, Jason D. Lee, Sham M. Kakade, Huan Wang, Caiming Xiong
摘要
Meta-learning aims to perform fast adaptation on a new task through learning a "prior" from multiple existing tasks. A common practice in meta-learning is to perform a train-validation split (train-val method ) where the prior adapts to the task on one split of the data, and the resulting predictor is evaluated on another split. Despite its prevalence, the importance of the train-validation split is not well understood either in theory or in practice, particularly in comparison to the more direct train-train method, which uses all the per-task data for both training and evaluation. We provide a detailed theoretical study on whether and when the train-validation split is helpful in the linear centroid meta-learning problem. In the agnostic case, we show that the expected loss of the train-val method is minimized at the optimal prior for meta testing, and this is not the case for the train-train method in general without structural assumptions on the data. In contrast, in the realizable case where the data are generated from linear models, we show that both the train-val and train-train losses are minimized at the optimal prior in expectation. Further, perhaps surprisingly, our main result shows that the train-train method achieves a strictly better excess loss in this realizable case, even when the regularization parameter and split ratio are optimally tuned for both methods. Our results highlight that sample splitting may not always be preferable, especially when the data is realizable by the model. We validate our theories by experimentally showing that the train-train method can indeed outperform the train-val method, on both simulations and real meta-learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- On Episodes, Prototypical Networks, and Few-Shot LearningSteinar Laenen, Luca BertinettoNeurIPS 2021 · 被引用 142 次
- Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective AdaptationHaoxiang Wang, Han Zhao, Bo LiICML 2021 · 被引用 108 次
- AutoBalance: Optimized Loss Functions for Imbalanced DataMingchen Li, Xuechen Zhang, Christos Thrampoulidis, Jiasi Chen 等NeurIPS 2021 · 被引用 89 次
- Generalization Bounds For Meta-Learning: An Information-Theoretic AnalysisQi Chen, Changjian Shui, Mario MarchandNeurIPS 2021 · 被引用 66 次
它引用的顶会 Paper10
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 被引用 218 次
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 被引用 128 次
- Convergence of Meta-Learning with Task-Specific Adaptation over Partial ParametersKaiyi Ji, Jason D. Lee, Yingbin Liang, H. Vincent PoorNeurIPS 2020 · 被引用 97 次
相关 Paper
- Understanding Train-Validation Split in Meta-Learning with Neural NetworksXinzhe Zuo, Zixiang Chen, Huaxiu Yao, Yuan Cao 等ICLR 2023
- A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-LearningNikunj Saunshi, Arushi Gupta, Wei HuICML 2021 · 被引用 19 次
- The Role of Deconfounding in Meta-learningYinjie Jiang, Zhengyu Chen, Kun Kuang, Luotian Yuan 等ICML 2022 · 被引用 15 次
- Shallow Bayesian Meta Learning for Real-World Few-Shot RecognitionXueting Zhang, Debin Meng, Henry Gouk, Timothy M. HospedalesICCV 2021 · 被引用 88 次
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
