The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model Pretraining
Ting-Rui Chiang, Dani Yogatama
Abstract
We analyze the masked language modeling pretraining objective function from the perspective of the Distributional Hypothesis. We investigate whether the better sample efficiency and the better generalization capability of models pretrained with masked language modeling can be attributed to the semantic similarity encoded in the pretraining data’s distributional property. Via a synthetic dataset, our analysis suggests that distributional property indeed leads to the better sample efficiency of pretrained masked language models, but does not fully explain the generalization capability. We also conduct an analysis over two real-world datasets and demonstrate that the distributional property does not explain the generalization ability of pretrained natural language models either. Our results illustrate our limited understanding of model pretraining and provide future research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 425 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang et al.EMNLP 2022 · 148 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
Related papers
- On Masked Pre-training and the Marginal LikelihoodPablo Moreno-Muñoz, Pol Garcia Recasens, Søren HaubergNeurIPS 2023 · 8 citations
- Towards Semantics-Enhanced Pre-Training: Can Lexicon Definitions Help Learning Sentence Meanings?Xuancheng Ren, Xu Sun, Houfeng Wang, Qun LiuAAAI 2021 · 5 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim et al.EMNLP 2022 · 8 citations
- Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationJaeseong Lee, Dohyeon Lee, Seung-won HwangAAAI 2023 · 1 citation
