SHOT: Suppressing the Hessian along the Optimization Trajectory for Gradient-Based Meta-Learning
Junhoo Lee, Jayeon Yoo, Nojun Kwak
摘要
In this paper, we hypothesize that gradient-based meta-learning (GBML) implicitly suppresses the Hessian along the optimization trajectory in the inner loop. Based on this hypothesis, we introduce an algorithm called SHOT (Suppressing the Hessian along the Optimization Trajectory) that minimizes the distance between the parameters of the target and reference models to suppress the Hessian in the inner loop. Despite dealing with high-order terms, SHOT does not increase the computational complexity of the baseline model much. It is agnostic to both the algorithm and architecture used in GBML, making it highly versatile and applicable to any GBML baseline. To validate the effectiveness of SHOT, we conduct empirical tests on standard few-shot learning tasks and qualitatively analyze its dynamics. We confirm our hypothesis empirically and demonstrate that SHOT outperforms the corresponding baseline. Code is available at: https://github.com/JunHoo-Lee/SHOT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Any-Way Meta LearningJunhoo Lee, Yearim Kim, Hyunho Lee, Nojun KwakAAAI 2024 · 被引用 2 次
- Memory-Reduced Meta-Learning with Guaranteed ConvergenceHonglin Yang, Ji Ma, Xiao YuAAAI 2025 · 被引用 1 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
相关 Paper
- On the Convergence Theory for Hessian-Free Bilevel AlgorithmsDaouda Sow, Kaiyi Ji, Yingbin LiangNeurIPS 2022 · 被引用 51 次
- Sharp-MAML: Sharpness-Aware Model-Agnostic Meta LearningMomin Abbas, Quan Xiao, Lisha Chen, Pin-Yu Chen 等ICML 2022 · 被引用 105 次
- MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningBaoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 等AAAI 2024 · 被引用 91 次
- Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective AdaptationHaoxiang Wang, Han Zhao, Bo LiICML 2021 · 被引用 108 次
- Meta-Learning with a Geometry-Adaptive PreconditionerSuhyun Kang, Duhun Hwang, Moonjung Eo, Taesup Kim 等CVPR 2023
