Meta-Learning with Self-Improving Momentum Target
Jihoon Tack, Jongjin Park, Hankook Lee, Jaeho Lee, Jinwoo Shin
摘要
The idea of using a separately trained target model (or teacher) to improve the performance of the student model has been increasingly popular in various machine learning domains, and meta-learning is no exception; a recent discovery shows that utilizing task-wise target models can significantly boost the generalization performance. However, obtaining a target model for each task can be highly expensive, especially when the number of tasks for meta-learning is large. To tackle this issue, we propose a simple yet effective method, coined Self-improving Momentum Target (SiMT). SiMT generates the target model by adapting from the temporal ensemble of the meta-learner, i.e., the momentum network. This momentum network and its task-specific adaptations enjoy a favorable generalization performance, enabling self-improving of the meta-learner through knowledge distillation. Moreover, we found that perturbing parameters of the meta-learner, e.g., dropout, further stabilize this self-improving process by preventing fast convergence of the distillation loss during meta-training. Our experimental results demonstrate that SiMT brings a significant performance gain when combined with a wide range of meta-learning methods under various applications, including few-shot regression, few-shot classification, and meta-reinforcement learning. Code is available at https://github.com/jihoontack/SiMT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Meta-AdaM: An Meta-Learned Adaptive Optimizer with Momentum for Few-Shot LearningSiyuan Sun, Hongyang GaoNeurIPS 2023 · 被引用 51 次
- Learning Large-scale Neural Fields via Context Pruned Meta-LearningJihoon Tack, Subin Kim, Sihyun Yu, Jaeho Lee 等NeurIPS 2023 · 被引用 16 次
- Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-EncoderHuiwon Jang, Jihoon Tack, Daewon Choi, Jongheon Jeong 等NeurIPS 2023 · 被引用 9 次
- STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled TablesJaehyun Nam, Jihoon Tack, Kyungmin Lee, Hankook Lee 等ICLR 2023 · 被引用 2 次
- Revisiting Neural Networks for Few-Shot Learning: A Zero-Cost NAS PerspectiveHaidong KangICML 2025
它引用的顶会 Paper21
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- Towards Enabling Meta-Learning from Target ModelsSu Lu, Han-Jia Ye, Le Gan, De-Chuan ZhanNeurIPS 2021 · 被引用 6 次
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 被引用 8 次
- Meta Dropout: Learning to Perturb Latent Features for GeneralizationHaebeom Lee, Taewook Nam, Eunho Yang, Sung Ju HwangICLR 2020 · 被引用 59 次
- MATE: Plugging in Model Awareness to Task Embedding for Meta LearningXiaohan Chen, Zhangyang Wang, Siyu Tang, Krikamol MuandetNeurIPS 2020 · 被引用 10 次
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 被引用 114 次
