Learning where to learn: Gradient sparsity in meta and continual learning
Johannes von Oswald, Dominic Zhao, Seijin Kobayashi, Simon Schug, Massimo Caccia, Nicolas Zucchet, João Sacramento
Abstract
Finding neural network weights that generalize well from small datasets is difficult. A promising approach is to learn a weight initialization such that a small number of weight changes results in low generalization error. We show that this form of meta-learning can be improved by letting the learning algorithm decide which weights to change, i.e., by learning where to learn. We find that patterned sparsity emerges from this process, with the pattern of sparsity varying on a problem-by-problem basis. This selective sparsity results in better generalization and less interference in a range of few-shot and continual learning problems. Moreover, we find that sparse learning also emerges in a more expressive model where learning rates are meta-learned. Our results shed light on an ongoing debate on whether meta-learning can discover adaptable features and suggest that learning by sparse gradient descent is a powerful inductive bias for meta-learning systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7b464f0-c821-4d93-ace1-fa7466bfa45fCited by top-tier papers25
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 140 citations
- Continual Learning via Local Module CompositionOleksiy Ostapenko, Pau Rodríguez, Massimo Caccia, Laurent CharlinNeurIPS 2021 · 98 citations
- MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningBaoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li et al.AAAI 2024 · 91 citations
Builds on6
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
- BOIL: Towards Representation Change for Few-shot LearningJaehoon Oh, Hyungjun Yoo, ChangHwan Kim, Se-Young YunICLR 2021 · 185 citations
- TaskNorm: Rethinking Batch Normalization for Meta-LearningJohn Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin et al.ICML 2020 · 93 citations
- Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningMassimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin et al.NeurIPS 2020 · 83 citations
- Modular Meta-Learning with ShrinkageYutian Chen, Abram L. Friesen, Feryal M. P. Behbahani, Arnaud Doucet et al.NeurIPS 2020 · 35 citations
Related papers
- Meta-ticket: Finding optimal subnetworks for few-shot learning within randomly initialized neural networksDaiki Chijiwa, Shin'ya Yamaguchi, Atsutoshi Kumagai, Yasutoshi IdaNeurIPS 2022 · 12 citations
- Meta-Learning with Adaptive HyperparametersSungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim et al.NeurIPS 2020 · 164 citations
- MetaNorm: Learning to Normalize Few-Shot Batches Across DomainsYing-Jun Du, Xiantong Zhen, Ling Shao, Cees G. M. SnoekICLR 2021 · 26 citations
- Large-Scale Meta-Learning with Continual Trajectory ShiftingJaewoong Shin, Haebeom Lee, Boqing Gong, Sung Ju HwangICML 2021 · 18 citations
- Towards Sample-efficient Overparameterized Meta-learningYue Sun, Adhyyan Narang, Halil Ibrahim Gulluk, Samet Oymak et al.NeurIPS 2021 · 26 citations
