Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza
Abstract
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data. However, the generalizability of these results across languages, model types, and evaluation settings remains unclear. We test this by comparing models trained on CDL vs. Wikipedia across two LM objectives (masked and causal), three languages (English, French, German), and three syntactic minimal-pair benchmarks. Our results on these benchmarks show inconsistent benefits of CDL, which in most cases is outperformed by Wikipedia models. We then identify various shortcomings in previous benchmarks, and introduce a novel testing methodology, FIT-CLAMS, which uses a frequency-controlled design to enable balanced comparisons across training corpora. Through minimal pair evaluations and regression analysis we show that training on CDL does not yield stronger generalizations for acquiring syntax and highlight the importance of controlling for frequency effects when evaluating syntactic ability. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9bfd0f3-f3d6-473f-ac26-e5934a0465b6Cited by top-tier papers1
Ask how each one uses itBuilds on3
- How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive BiasesAaron Mueller, Tal LinzenACL 2023 · 9 citations
- Cross-Linguistic Syntactic Evaluation of Word Prediction ModelsAaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina et al.ACL 2020 · 2 citations
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
Related papers
- Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge InjectionYuwei Zhang, Wenhao Yu, Shangbin Feng, Yifan Zhu et al.ACL 2026 · 7 citations
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 17 citations
- Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial LanguagesNadine El-Naggar, Tatsuki Kuribayashi, Ted BriscoeEMNLP 2025
- Understanding Data Temporality Impact on Large Language Models Pre-trainingRomain Fabre, Hippolyte Pilchen, Franck SIGNE TALLA, Patrick Perez et al.ICML 2026
- BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant LearningShengao Wang, Arjun Chandra, Aoming Liu, Venkatesh Saligrama et al.ICCV 2025 · 8 citations
