Unraveling Syntax: Language Modeling and the Substructure of Grammars
Laura Ying Schulz, Daniel Mitropolsky, Tomaso A Poggio
Abstract
While large models achieve impressive results, their learning dynamics are far from understood. Many domains of interest -- such as natural language syntax, coding languages, arithmetic problems -- are captured by context-free grammars (CFGs). In this work, we extend prior work on neural language modeling of CFGs in a novel direction: how language modeling behaves with respect to CFG substructure, namely "subgrammars". We first define subgrammars, and prove a set of fundamental theorems regarding language modeling and subgrammars. We show that language modeling loss (or equivalently the Kullback-Leibler divergence) recurses linearly over its top-level subgrammars; applied recursively, the loss decomposes into losses for "irreducible" subgrammars. We also prove that the constant in this linear recurrence is a function of the expected "recursion", a notion we introduce. We show that under additional assumptions, parametrized models learn subgrammars in parallel. Empirically, we confirm that small transformers learn subgrammars in parallel, unlike children, who first master simple substructures. We also briefly explore several other questions regarding subgrammars. We find that subgrammar pretraining can improve final performance, but only for tiny models relative to the grammar, while alignment analyses show that pretraining consistently lead to internal representations that better reflect the grammar’s substructure in all cases; we also observe persistent difficulty with deeper recursion, a limitation that appears even of large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e312180-1dd3-430f-9482-d2a57bf57266Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 104 citations
- Neural Networks and the Chomsky HierarchyGrégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein et al.ICLR 2023 · 45 citations
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language ModelsIsabel Papadimitriou, Dan JurafskyEMNLP 2020 · 40 citations
Related papers
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 4 citations
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 33 citations
- Context-free Recognition with TransformersSelim Jerad, Anej Svete, Sophie Hao, Ryan Cotterell et al.ICML 2026 · 3 citations
- Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar InductionJinwook Park, Kangil KimEMNLP 2025
- In-Context Language Learning: Architectures and AlgorithmsEkin Akyürek, Bailin Wang, Yoon Kim, Jacob AndreasICML 2024 · 91 citations
