Efficient Variance Reduction for Meta-learning
Hansi Yang, James T. Kwok
Abstract
Meta-learning tries to learn meta-knowledge from a large number of tasks. However, the stochastic meta-gradient can have large variance due to data sampling (from each task) and task sampling (from the whole task distribution), leading to slow convergence. In this paper, we propose a novel approach that integrates variance reduction with first-order meta-learning algorithms such as Reptile. It retains the bilevel formulation which better captures the structure of meta-learning, but does not require storing the vast number of taskspecific parameters in general bilevel variance reduction methods. Theoretical results show that it has fast convergence rate due to variance reduction. Experiments on benchmark few-shot classification data sets demonstrate its effectiveness over state-of-the-art meta-learning algorithms with and without variance reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Meta Continual Learning Revisited: Implicitly Enhancing Online Hessian Approximation via Variance ReductionYichen Wu, Long-Kai Huang, Renzhen Wang, Deyu Meng et al.ICLR 2024 · 42 citations
- MetaFBP: Learning to Learn High-Order Predictor for Personalized Facial Beauty PredictionLuojun Lin, Zhifeng Shen, Jia-Li Yin, Qipeng Liu et al.ACM MM 2023 · 5 citations
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningMinyoung Kim, Timothy M. HospedalesAAAI 2025 · 3 citations
Builds on10
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 176 citations
- A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-MomentumPrashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai et al.NeurIPS 2021 · 175 citations
Related papers
- The Effect of Diversity in Meta-LearningRamnath Kumar, Tristan Deleu, Yoshua BengioAAAI 2023 · 18 citations
- On the Convergence Theory for Hessian-Free Bilevel AlgorithmsDaouda Sow, Kaiyi Ji, Yingbin LiangNeurIPS 2022 · 51 citations
- Meta-Learning of Neural Architectures for Few-Shot LearningThomas Elsken, Benedikt Staffler, Jan Hendrik Metzen, Frank HutterCVPR 2020
- Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective AdaptationHaoxiang Wang, Han Zhao, Bo LiICML 2021 · 108 citations
- Robust Meta-learning with Sampling Noise and Label Noise via Eigen-ReptileDong Chen, Lingfei Wu, Siliang Tang, Xiao Yun et al.ICML 2022 · 13 citations
