MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning
Baoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li, Huiwei Lin, Yunming Ye, Bowen Zhang
Abstract
Equipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the outer-loop process learns a shared gradient descent algorithm (called meta-optimizer), while the inner-loop process leverages it to optimize a task-specific base learner with few examples. Although these methods have shown superior performance on FSL, the outer-loop process requires calculating second-order derivatives along the inner-loop path, which imposes considerable memory burdens and the risk of vanishing gradients. This degrades meta-learning performance. Inspired by recent diffusion models, we find that the inner-loop gradient descent process can be viewed as a reverse process (i.e., denoising) of diffusion where the target of denoising is the weight of base learner but origin data. Based on this fact, we propose to model the gradient descent algorithm as a diffusion model and then present a novel conditional diffusion-based meta-learning, called MetaDiff, that effectively models the optimization process of base learner weights from Gaussian initialization to target weights in a denoising manner. Thanks to the training efficiency of diffusion models, our MetaDiff does not need to differentiate through the inner-loop path such that the memory burdens and the risk of vanishing gradients can be effectively alleviated for improving FSL. Experimental results show that our MetaDiff outperforms state-of-the-art gradient-based meta-learning family on FSL tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted DiffusionYongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang et al.NeurIPS 2024 · 15 citations
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu et al.NeurIPS 2025 · 10 citations
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li et al.CVPR 2026 · 6 citations
- Weight Diffusion for Future: Learn to Generalize in Non-Stationary EnvironmentsMixue Xie, Shuang Li, Binhui Xie, Chi Harold Liu et al.NeurIPS 2024 · 6 citations
- DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network WeightsSaumya Gupta, Scott Biggs, Moritz Laber, Zohair Shafi et al.ICLR 2026 · 5 citations
Builds on18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot LearningYinbo Chen, Zhuang Liu, Huijuan Xu, Trevor Darrell et al.ICCV 2021 · 455 citations
Related papers
- ProtoDiff: Learning to Learn Prototypical Networks by Task-Guided DiffusionYingjun Du, Zehao Xiao, Shengcai Liao, Cees SnoekNeurIPS 2023 · 33 citations
- MetaNODE: Prototype Optimization as a Neural ODE for Few-Shot LearningBaoquan Zhang, Xutao Li, Shanshan Feng, Yunming Ye et al.AAAI 2022 · 46 citations
- Meta-AdaM: An Meta-Learned Adaptive Optimizer with Momentum for Few-Shot LearningSiyuan Sun, Hongyang GaoNeurIPS 2023 · 51 citations
- Learning to Learn Weight Generation via Local Consistency DiffusionYunchuan Guan, Yu Liu, Ke Zhou, Zhiqi Shen et al.CVPR 2026 · 5 citations
- SHOT: Suppressing the Hessian along the Optimization Trajectory for Gradient-Based Meta-LearningJunhoo Lee, Jayeon Yoo, Nojun KwakNeurIPS 2023 · 4 citations
