How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speech
Aditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoy
Abstract
When acquiring syntax, children consistently choose hierarchical rules over competing non-hierarchical possibilities. Is this preference due to a learning bias for hierarchical structure, or due to more general biases that interact with hierarchical cues in children's linguistic input? We explore these possibilities by training LSTMs and Transformers - two types of neural networks without a hierarchical bias - on data similar in quantity and content to children's linguistic input: text from the CHILDES corpus. We then evaluate what these models have learned about English yes/no questions, a phenomenon for which hierarchical structure is crucial. We find that, though they perform well at capturing the surface statistics of child-directed speech (as measured by perplexity), both model types generalize in a way more consistent with an incorrect linear rule than the correct hierarchical rule. These results suggest that human-like generalization from text alone requires stronger biases than the general sequence-processing biases of standard neural network architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe93bea4-78ec-4a68-b370-5ce791afee58Cited by top-tier papers2
- Information Locality as an Inductive Bias for Neural Language ModelsTaiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell et al.ACL 2025 · 6 citations
- Mind the Gap: How BabyLMs Learn Filler-Gap DependenciesChi-Yun Chang, Xueyang Huang, Humaira Nasir, Shane Storks et al.EMNLP 2025
Builds on2
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu et al.EMNLP 2020 · 7 citations
Related papers
- Semantic Training Signals Promote Hierarchical Syntactic Generalization in TransformersAditya Yedetore, Najoung KimEMNLP 2024 · 4 citations
- How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive BiasesAaron Mueller, Tal LinzenACL 2023 · 9 citations
- Examining the Inductive Bias of Neural Language Models with Artificial LanguagesJennifer C. White, Ryan CotterellACL 2021
- Characterizing intrinsic compositionality in transformers with Tree ProjectionsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningICLR 2023 · 13 citations
- The Grammar-Learning Trajectories of Neural Language ModelsLeshem Choshen, Guy Hacohen, Daphna Weinshall, Omri AbendACL 2022
