Meta-Learning with Warped Gradient Descent
Sebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin, Hujun Yin, Raia Hadsell
摘要
Learning an efficient update rule from data that promotes rapid learning of new tasks from the same distribution remains an open problem in meta-learning. Typically, previous works have approached this issue either by attempting to train a neural network that directly produces updates or by attempting to learn better initialisations or scaling factors for a gradient-based update rule. Both of these approaches pose challenges. On one hand, directly producing an update forgoes a useful inductive bias and can easily lead to non-converging behaviour. On the other hand, approaches that try to control a gradient-based update rule typically resort to computing gradients through the learning process to obtain their meta-gradients, leading to methods that can not scale beyond few-shot task adaptation. In this work, we propose Warped Gradient Descent (WarpGrad), a method that intersects these approaches to mitigate their limitations. WarpGrad meta-learns an efficiently parameterised preconditioning matrix that facilitates gradient descent across the task distribution. Preconditioning arises by interleaving non-linear layers, referred to as warp-layers, between the layers of a task-learner. Warp-layers are meta-learned without backpropagating through the task training process in a manner similar to methods that learn to directly produce updates. WarpGrad is computationally efficient, easy to implement, and can scale to arbitrarily large meta-learning problems. We provide a geometrical interpretation of the approach and evaluate its effectiveness in a variety of settings, including few-shot, standard supervised, continual and reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper71
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- FedBABU: Toward Enhanced Representation for Federated Image ClassificationJaehoon Oh, Sangmook Kim, Se-Young YunICLR 2022 · 被引用 332 次
- BOIL: Towards Representation Change for Few-shot LearningJaehoon Oh, Hyungjun Yoo, ChangHwan Kim, Se-Young YunICLR 2021 · 被引用 185 次
- Meta-Learning with Adaptive HyperparametersSungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim 等NeurIPS 2020 · 被引用 164 次
相关 Paper
- On Enforcing Better Conditioned Meta-Learning for Rapid Few-Shot AdaptationMarkus Hiller, Mehrtash Harandi, Tom DrummondNeurIPS 2022 · 被引用 10 次
- Meta-Learning with a Geometry-Adaptive PreconditionerSuhyun Kang, Duhun Hwang, Moonjung Eo, Taesup Kim 等CVPR 2023
- Learning where to learn: Gradient sparsity in meta and continual learningJohannes von Oswald, Dominic Zhao, Seijin Kobayashi, Simon Schug 等NeurIPS 2021 · 被引用 61 次
- MetaFun: Meta-Learning with Iterative Functional UpdatesJin Xu, Jean-Francois Ton, Hyunjik Kim, Adam R. Kosiorek 等ICML 2020 · 被引用 76 次
- Large-Scale Meta-Learning with Continual Trajectory ShiftingJaewoong Shin, Haebeom Lee, Boqing Gong, Sung Ju HwangICML 2021 · 被引用 18 次
