Function Words as Statistical Cues for Language Learning
Xiulin Yang, Heidi R. Getz, Ethan Gotlieb Wilcox
Abstract
What statistical properties might support learning abstract grammatical knowledge from linear input? We address this question by examining the statistical distribution of function words. Function words have been argued to aid acquisition through three distributional properties: high frequency, reliable syntactic association, and phrase-boundary alignment. We conduct a cross-linguistic corpus analysis of 186 languages, which confirms that all three properties are universal. Using counterfactual language modeling and ablation experiments on English, we show that preserving these properties facilitates acquisition in neural learners, with a Goldilocks effect: function words must be frequent enough to be reliable, yet diverse enough to remain informative to structural dependency. Probing analyses further reveal that different learning conditions produce systematically different reliance on function words. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 447746e4-9566-4c4e-9c78-cd117bc79de9Builds on5
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald et al.ACL 2024 · 15 citations
- Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNsKanishka Misra, Kyle MahowaldEMNLP 2024 · 11 citations
- Language Models Grow Less Humanlike beyond Phase TransitionTatsuya Aoyama, Ethan WilcoxACL 2025
- Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not MeaningWesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan SchneiderEMNLP 2025
Related papers
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 6 citations
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita et al.EMNLP 2020 · 2 citations
- Constructions are Revealed in Word DistributionsJoshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory ShainEMNLP 2025 · 8 citations
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 40 citations
- Pretraining with Artificial Language: Studying Transferable Knowledge in Language ModelsRyokan Ri, Yoshimasa TsuruokaACL 2022 · 40 citations
