ICLR2024
A Hierarchical Bayesian Model for Few-Shot Meta Learning
Minyoung Kim, Timothy M. Hospedales
被引用 7 次
摘要
We propose novel model parametrisation and inference algorithm in a hierarchical Bayesian model for the few-shot meta learning problem. We consider episode-wise random variables to model episode-specific generative processes, where these local random variables are governed by a higher-level global random variable. The global variable captures information shared across episodes, while controlling how much the model needs to be adapted to new episodes in a principled Bayesian manner. Within our framework, prediction on a novel episode/task can be seen as a Bayesian inference problem. For tractable training, we need to be able to relate each local episode-specific solution to the global higher-level parameters. We propose a Normal-Inverse-Wishart model, for which establishing this localglobal relationship becomes feasible due to the approximate closed-form solutions for the local posterior distributions. The resulting algorithm is more attractive than the MAML in that it does not maintain a costly computational graph for the sequence of gradient descent steps in an episode. Our approach is also different from existing Bayesian meta learning methods in that rather than modeling a single random variable for all episodes, it leverages a hierarchical structure that exploits the local-global relationships desirable for principled Bayesian learning with many related tasks. where the variational parameters L consists of L 0 (parameters for q(ϕ)) and L i N i=1 's (parameters of q i (θ i )'s for episode i). Note that although θ i 's are independent across episodes under (2), they The second term of (95) becomes a Gaussian following the derivation similar to (88) with the test support data X * and Y * included. Consequently we let p(ϕ|D * , D 1 , . . . , D N ) = N (ϕ; A -1 * b * , A -1 * ). At last, (95) is the marginalisation of the product of two Gaussians, which admits the following closed form: