Online Hyperparameter Meta-Learning with Hypergradient Distillation
Haebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang, Timothy M. Hospedales, Sung Ju Hwang
摘要
Many gradient-based meta-learning methods assume a set of parameters that do not participate in inner-optimization, which can be considered as hyperparameters. Although such hyperparameters can be optimized using the existing gradient-based hyperparameter optimization (HO) methods, they suffer from the following issues. Unrolled differentiation methods do not scale well to high-dimensional hyperparameters or horizon length, Implicit Function Theorem (IFT) based methods are restrictive for online optimization, and short horizon approximations suffer from short horizon bias. In this work, we propose a novel HO method that can overcome these limitations, by approximating the second-order term with knowledge distillation. Specifically, we parameterize a single Jacobian-vector product (JVP) for each HO step and minimize the distance from the true second-order term. Our method allows online optimization and also is scalable to the hyperparameter dimension and the horizon length. We demonstrate the effectiveness of our method on two different meta-learning methods and three benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Cross-Domain Few-Shot Classification via Learned Feature-Wise TransformationHung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, Ming-Hsuan YangICLR 2020 · 被引用 467 次
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin 等ICLR 2020 · 被引用 221 次
- Meta Dropout: Learning to Perturb Latent Features for GeneralizationHaebeom Lee, Taewook Nam, Eunho Yang, Sung Ju HwangICLR 2020 · 被引用 59 次
- Large-Scale Meta-Learning with Continual Trajectory ShiftingJaewoong Shin, Haebeom Lee, Boqing Gong, Sung Ju HwangICML 2021 · 被引用 18 次
- MetaPerturb: Transferable Regularizer for Heterogeneous Tasks and ArchitecturesJeongun Ryu, Jaewoong Shin, Haebeom Lee, Sung Ju HwangNeurIPS 2020 · 被引用 8 次
相关 Paper
- The Curse of Unrolling: Rate of Differentiating Through OptimizationDamien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian PedregosaNeurIPS 2022 · 被引用 20 次
- On Implicit Bias in Overparameterized Bilevel OptimizationPaul Vicol, Jonathan P. Lorraine, Fabian Pedregosa, David Duvenaud 等ICML 2022 · 被引用 48 次
- On the Convergence Theory for Hessian-Free Bilevel AlgorithmsDaouda Sow, Kaiyi Ji, Yingbin LiangNeurIPS 2022 · 被引用 51 次
- Gradient-based Hyperparameter Optimization Over Long HorizonsPaul Micaelli, Amos J. StorkeyNeurIPS 2021 · 被引用 23 次
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningMinyoung Kim, Timothy M. HospedalesAAAI 2025 · 被引用 3 次
