Meta-Learning without Memorization
Mingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine, Chelsea Finn
Abstract
The ability to learn new concepts with small amounts of data is a critical aspect of intelligence that has proven challenging for deep learning methods. Meta-learning has emerged as a promising technique for leveraging data from previous tasks to enable efficient learning of new tasks. However, most meta-learning algorithms implicitly require that the meta-training tasks be mutually-exclusive, such that no single model can solve all of the tasks at once. For example, when creating tasks for few-shot image classification, prior work uses a per-task random assignment of image classes to N-way classification labels. If this is not done, the meta-learner can ignore the task training data and learn a single model that performs all of the meta-training tasks zero-shot, but does not adapt effectively to new image classes. This requirement means that the user must take great care in designing the tasks, for example by shuffling labels or removing task identifying information from the inputs. In some domains, this makes meta-learning entirely inapplicable. In this paper, we address this challenge by designing a meta-regularization objective using information theory that places precedence on data-driven adaptation. This causes the meta-learner to decide what must be learned from the task training data and what should be inferred from the task testing input. By doing so, our algorithm can successfully use data from non-mutually-exclusive tasks to efficiently adapt to novel tasks. We demonstrate its applicability to both contextual and gradient-based meta-learning algorithms, and apply it in practical settings where applying standard meta-learning has been difficult. Our approach substantially outperforms standard meta-learning algorithms in these settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76c02851-7620-4c08-a7e2-f23cd1b2c2beCited by top-tier papers57
- BOIL: Towards Representation Change for Few-shot LearningJaehoon Oh, Hyungjun Yoo, ChangHwan Kim, Se-Young YunICLR 2021 · 185 citations
- Reducing Information Bottleneck for Weakly Supervised Semantic SegmentationJungbeom Lee, Jooyoung Choi, Jisoo Mok, Sungroh YoonNeurIPS 2021 · 174 citations
- Meta-Learning with Adaptive HyperparametersSungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim et al.NeurIPS 2020 · 164 citations
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos et al.ICML 2020 · 153 citations
- PACOH: Bayes-Optimal Meta-Learning with PAC-GuaranteesJonas Rothfuss, Vincent Fortuin, Martin Josifoski, Andreas KrauseICML 2021 · 136 citations
Related papers
- Data Augmentation for Meta-LearningRenkun Ni, Micah Goldblum, Amr Sharaf, Kezhi Kong et al.ICML 2021 · 93 citations
- Theoretical bounds on estimation error for meta-learningJames Lucas, Mengye Ren, Irene Raissa Kameni, Toniann Pitassi et al.ICLR 2021 · 12 citations
- Boosting Few-Shot Learning With Adaptive Margin LossAoxue Li, Weiran Huang, Xu Lan, Jiashi Feng et al.CVPR 2020
- Semantic matching for text classification with complex class descriptionsBrian de Silva, Kuan-Wen Huang, Gwang Lee, Karen Hovsepian et al.EMNLP 2023 · 1 citation
- Attentive Weights Generation for Few Shot Learning via Information MaximizationYiluan Guo, Ngai-Man CheungCVPR 2020
