Deep networks learn to parse uniform-depth context-free languages from local statistics
Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
Abstract
Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make these feats possible, and how much data is required, remain largely unknown. Probabilistic context-free grammars (PCFGs) provide a tractable testbed for studying these questions. However, prior work has focused either on the post-hoc characterization of the parsing-like algorithms used by trained networks; or on the learnability of PCFGs with fixed syntax, where parsing is unnecessary. Here, we (i) introduce a tunable class of PCFGs in which both the degree of ambiguity and the correlation structure across scales can be controlled; (ii) provide a learning mechanism---an inference algorithm inspired by the structure of deep convolutional networks---that links learnability and sample complexity to specific language statistics; and (iii) validate our predictions empirically across deep convolutional and transformer-based architectures. Overall, we propose a unifying framework where correlations at different scales lift local ambiguities, enabling the emergence of hierarchical representations of the data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3285ade3-14d0-46de-9557-a0aa1e964089Builds on5
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 33 citations
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz et al.NeurIPS 2024 · 27 citations
- Do Transformers Parse while Predicting the Masked Word?Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev AroraEMNLP 2023 · 5 citations
- How Compositional Generalization and Creativity Improve as Diffusion Models are TrainedAlessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard et al.ICML 2025
Related papers
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Unraveling Syntax: Language Modeling and the Substructure of GrammarsLaura Ying Schulz, Daniel Mitropolsky, Tomaso A PoggioICML 2026 · 5 citations
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda et al.ACL 2024
- Learning curves theory for hierarchically compositional data with power-law distributed featuresFrancesco Cagnetta, Hyunmo Kang, Matthieu WyartICML 2025
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 9 citations
