A Hierarchical Bayesian Model for Few-Shot Meta Learning
Minyoung Kim, Timothy M. Hospedales
摘要
We propose novel model parametrisation and inference algorithm in a hierarchical Bayesian model for the few-shot meta learning problem. We consider episode-wise random variables to model episode-specific generative processes, where these local random variables are governed by a higher-level global random variable. The global variable captures information shared across episodes, while controlling how much the model needs to be adapted to new episodes in a principled Bayesian manner. Within our framework, prediction on a novel episode/task can be seen as a Bayesian inference problem. For tractable training, we need to be able to relate each local episode-specific solution to the global higher-level parameters. We propose a Normal-Inverse-Wishart model, for which establishing this localglobal relationship becomes feasible due to the approximate closed-form solutions for the local posterior distributions. The resulting algorithm is more attractive than the MAML in that it does not maintain a costly computational graph for the sequence of gradient descent steps in an episode. Our approach is also different from existing Bayesian meta learning methods in that rather than modeling a single random variable for all episodes, it leverages a hierarchical structure that exploits the local-global relationships desirable for principled Bayesian learning with many related tasks. where the variational parameters L consists of L 0 (parameters for q(ϕ)) and L i N i=1 's (parameters of q i (θ i )'s for episode i). Note that although θ i 's are independent across episodes under (2), they The second term of (95) becomes a Gaussian following the derivation similar to (88) with the test support data X * and Y * included. Consequently we let p(ϕ|D * , D 1 , . . . , D N ) = N (ϕ; A -1 * b * , A -1 * ). At last, (95) is the marginalisation of the product of two Gaussians, which admits the following closed form:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu 等NeurIPS 2025 · 被引用 10 次
- DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot LearningWenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han 等ICLR 2026 · 被引用 4 次
- Identity-Free Deferral For Unseen ExpertsJoshua Strong, Pramit Saha, Yasin Ibrahim, Cheng Ouyang 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Meta-Learning without MemorizationMingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine 等ICLR 2020 · 被引用 201 次
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle 等NeurIPS 2020 · 被引用 167 次
- Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a DifferenceShell Xu Hu, Da Li, Jan Stühmer, Minyoung Kim 等CVPR 2022 · 被引用 161 次
相关 Paper
- Addressing Catastrophic Forgetting in Few-Shot ProblemsPau Ching Yap, Hippolyt Ritter, David BarberICML 2021 · 被引用 20 次
- Meta-GMVAE: Mixture of Gaussian VAE for Unsupervised Meta-LearningDong Bok Lee, Dongchan Min, Seanie Lee, Sung Ju HwangICLR 2021 · 被引用 62 次
- Few-shot Relation Extraction via Bayesian Meta-learning on Relation GraphsMeng Qu, Tianyu Gao, Louis-Pascal A. C. Xhonneux, Jian TangICML 2020 · 被引用 131 次
- Few-Shot Learning With Global Class RepresentationsAoxue Li, Tiange Luo, Tao Xiang, Weiran Huang 等ICCV 2019 · 被引用 119 次
- Learning to Balance: Bayesian Meta-Learning for Imbalanced and Out-of-distribution TasksHaebeom Lee, Hayeon Lee, Donghyun Na, Saehoon Kim 等ICLR 2020 · 被引用 115 次
