Semantic Training Signals Promote Hierarchical Syntactic Generalization in Transformers
Aditya Yedetore, Najoung Kim
Abstract
Neural networks without hierarchical biases often struggle to learn linguistic rules that come naturally to humans. However, neural networks are trained primarily on form alone, while children acquiring language additionally receive data about meaning. Would neural networks generalize more like humans when trained on both form and meaning? We investigate this by examining if Transformers—neural networks without a hierarchical bias—better achieve hierarchical generalization when trained on both form and meaning compared to when trained on form alone. Our results show that Transformers trained on form and meaning do favor the hierarchical generalization more than those trained on form alone, suggesting that statistical learners without hierarchical biases can leverage semantic training signals to bootstrap hierarchical syntactic generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b55da2f3-dc1a-47e4-bbd4-c35c451b08dfBuilds on2
Related papers
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 17 citations
- How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive BiasesAaron Mueller, Tal LinzenACL 2023 · 9 citations
- Characterizing intrinsic compositionality in transformers with Tree ProjectionsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningICLR 2023 · 13 citations
- Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional EncodingDaichi Hayakawa, Issei SatoACL 2025
- Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not MeaningWesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan SchneiderEMNLP 2025
