Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
Kanishka Misra, Kyle Mahowald
Abstract
Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question. To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction ("a beautiful five days"). We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed. We found that AANNs were still learned better than systematically perturbed variants of the construction. Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., "a few days"). An additional experiment showed that this learning is enhanced when there is more variability in the input. Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena. Data and code: https:// github.com/kanishkamisra/aannalysis .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of RaceLihao Sun, Chengzhi Mao, Valentin Hofmann, Xuechunzi BaiACL 2025 · 13 citations
- Constructions are Revealed in Word DistributionsJoshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory ShainEMNLP 2025 · 8 citations
- Function Words as Statistical Cues for Language LearningXiulin Yang, Heidi R. Getz, Ethan Gotlieb WilcoxACL 2026 · 1 citation
- On the Acquisition of Shared Grammatical Representations in Bilingual Language ModelsCatherine Arnett, Tyler A. Chang, James A. Michaelov, Ben BergenACL 2025
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi et al.ACL 2026
Builds on5
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 40 citations
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz et al.ACL 2022 · 38 citations
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald et al.ACL 2024 · 15 citations
- The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative CorrelativeLeonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich SchützeEMNLP 2022 · 13 citations
Related papers
- Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not MeaningWesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan SchneiderEMNLP 2025
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita et al.EMNLP 2020 · 2 citations
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 6 citations
- On the Ability and Limitations of Transformers to Recognize Formal LanguagesSatwik Bhattamishra, Kabir Ahuja, Navin GoyalEMNLP 2020 · 7 citations
- A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding SpacesGabriella Chronis, Kyle Mahowald, Katrin ErkACL 2023 · 3 citations
