Theoretical Characterization of the Generalization Performance of Overfitted Meta-Learning
Peizhong Ju, Yingbin Liang, Ness B. Shroff
Abstract
Meta-learning has arisen as a successful method for improving training performance by training over many similar tasks, especially with deep neural networks (DNNs). However, the theoretical understanding of when and why overparameterized models such as DNNs can generalize well in meta-learning is still limited. As an initial step towards addressing this challenge, this paper studies the generalization performance of overfitted meta-learning under a linear regression model with Gaussian features. In contrast to a few recent studies along the same line, our framework allows the number of model parameters to be arbitrarily larger than the number of features in the ground truth signal, and hence naturally captures the overparameterized regime in practical deep meta-learning. We show that the overfitted min -norm solution of model-agnostic meta-learning (MAML) can be beneficial, which is similar to the recent remarkable findings on benign overfitting'' and double descent'' phenomenon in the classical (single-task) linear regression. However, due to the uniqueness of meta-learning such as task-specific gradient descent inner training and the diversity/fluctuation of the ground-truth signals among training tasks, we find new and interesting properties that do not exist in single-task linear regression. We first provide a high-probability upper bound (under reasonable tightness) on the generalization error, where certain terms decrease when the number of features increases. Our analysis suggests that benign overfitting is more significant and easier to observe when the noise and the diversity/fluctuation of the ground truth of each training task are large. Under this circumstance, we show that the overfitted min -norm solution can achieve an even lower generalization error than the underparameterized solution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25da3570-b7e1-4fda-806b-6a6791a4c318Cited by top-tier papers2
- Transfer Learning for Benign Overfitting in High-Dimensional Linear RegressionYeichan Kim, Ilmun Kim, Seyoung ParkNeurIPS 2025 · 2 citations
- Unlocking the Power of Rehearsal in Continual Learning: A Theoretical PerspectiveJunze Deng, Qinhang Wu, Peizhong Ju, Sen Lin et al.ICML 2025
Builds on9
- How Important is the Train-Validation Split in Meta-Learning?Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao et al.ICML 2021 · 60 citations
- Towards Sample-efficient Overparameterized Meta-learningYue Sun, Adhyyan Narang, Halil Ibrahim Gulluk, Samet Oymak et al.NeurIPS 2021 · 26 citations
- Overfitting Can Be Harmless for Basis Pursuit, But Only to a DegreePeizhong Ju, Xiaojun Lin, Jia LiuNeurIPS 2020 · 20 citations
- A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-LearningNikunj Saunshi, Arushi Gupta, Wei HuICML 2021 · 19 citations
- Provable Generalization of Overparameterized Meta-learning Trained with SGDYu Huang, Yingbin Liang, Longbo HuangNeurIPS 2022 · 14 citations
Related papers
- Understanding Benign Overfitting in Gradient-Based Meta LearningLisha Chen, Songtao Lu, Tianyi ChenNeurIPS 2022 · 20 citations
- Meta-learning with negative learning ratesAlberto BernacchiaICLR 2021 · 4 citations
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan et al.ICML 2020 · 40 citations
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki et al.NeurIPS 2021 · 12 citations
- On the Generalization Power of Overfitted Two-Layer Neural Tangent Kernel ModelsPeizhong Ju, Xiaojun Lin, Ness B. ShroffICML 2021 · 13 citations
