Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, Douwe Kiela
Abstract
A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines. In this paper, we propose a different explanation: MLMs succeed on downstream tasks mostly due to their ability to model higher-order word cooccurrence statistics. To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and we show that these models still achieve high accuracy after finetuning on many downstream tasks -including tasks specifically designed to be challenging for models that ignore word order. Our models also perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information. Overall, our results show that purely distributional information largely explains the success of pretraining, and they underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 311241ea-b4c1-432f-a940-77264a2d8e6cCited by top-tier papers43
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise ExpressionsYunlong Wang, Shuyuan Shen, Brian Y. LimCHI 2023 · 118 citations
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau et al.ACL 2022 · 72 citations
Builds on14
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 384 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 163 citations
- Efficient Second-Order TreeCRF for Neural Dependency ParsingYu Zhang, Zhenghua Li, Min ZhangACL 2020 · 90 citations
Related papers
- Word Order Does Matter and Shuffled Language Models Know ItMostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders SøgaardACL 2022
- On the Transferability of Pre-trained Language Models: A Study from Artificial DatasetsDavid Cheng-Han Chiang, Hung-Yi LeeAAAI 2022 · 33 citations
- Cross-Lingual Ability of Multilingual Masked Language Models: A Study of Language StructureYuan Chai, Yaobo Liang, Nan DuanACL 2022
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend et al.ICLR 2021 · 83 citations
- Pseudo Zero Pronoun Resolution Improves Zero Anaphora ResolutionRyuto Konno, Shun Kiyono, Yuichiroh Matsubayashi, Hiroki Ouchi et al.EMNLP 2021 · 6 citations
