Frequency Effects on Syntactic Rule Learning in Transformers
Jason Wei, Dan Garrette, Tal Linzen, Ellie Pavlick
Abstract
Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules. We investigate this question using the case study of BERT's performance on English subject-verb agreement. Unlike prior work, we train multiple instances of BERT from scratch, allowing us to perform a series of controlled interventions at pre-training time. We show that BERT often generalizes well to subject-verb pairs that never occurred in training, suggesting a degree of rule-governed behavior. We also find, however, that performance is heavily influenced by word frequency, with experiments showing that both the absolute frequency of a verb form, as well as the frequency relative to the alternate inflection, are causally implicated in the predictions BERT makes at inference time. Closer analysis of these frequency effects reveals that BERT's behavior is consistent with a system that correctly applies the SVA rule in general but struggles to overcome strong training priors and to estimate agreement features (singular vs. plural) on infrequent lexical items. * Work done while visiting Google.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f02d0d3-fb20-416e-9599-a93da05b66ddCited by top-tier papers13
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz et al.ACL 2022 · 38 citations
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and ScaleJames A. Michaelov, Roger P. Levy, Benjamin BergenNeurIPS 2025 · 15 citations
- The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative CorrelativeLeonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich SchützeEMNLP 2022 · 13 citations
- Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNsKanishka Misra, Kyle MahowaldEMNLP 2024 · 11 citations
Builds on6
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Predicting Inductive Biases of Pre-Trained ModelsCharles Lovering, Rohan Jha, Tal Linzen, Ellie PavlickICLR 2021 · 70 citations
- Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu et al.EMNLP 2020 · 7 citations
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 6 citations
Related papers
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau et al.ACL 2022 · 72 citations
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber et al.ACL 2021
- Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic SmoothingRichard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery et al.EMNLP 2024 · 1 citation
- BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model PerformanceTimo Schick, Hinrich SchützeACL 2020 · 1 citation
- MTLS: Making Texts into Linguistic SymbolsWenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang et al.EMNLP 2024 · 1 citation
