Lune

NeurIPS2023Top-tier venue

On Masked Pre-training and the Marginal Likelihood

Pablo Moreno-Muñoz, Pol Garcia Recasens, Søren Hauberg

2023Year
8Citations

Abstract

Masked pre-training removes random input dimensions and learns a model that can predict the missing values. Empirical results indicate that this intuitive form of self-supervised learning yields models that generalize very well to new domains. A theoretical understanding is, however, lacking. This paper shows that masked pre-training with a suitable cumulative scoring function corresponds to maximizing the model's marginal likelihood, which is de facto the Bayesian model selection measure of generalization. Beyond shedding light on the success of masked pre-training, this insight also suggests that Bayesian models can be trained with appropriately designed self-supervision. Empirically, we confirm the developed theory and explore the main learning principles of masked pre-training in large language models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cd4c0d4e-952b-46f4-80f9-6ae3ad4c7a2b

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines