Data Drives Unstable Hierarchical Generalization in LMs
Tian Qin, Naomi Saphra, David Alvarez-Melis
Abstract
Early in training, LMs can behave like n-gram models, but eventually, they often learn treebased syntactic rules and generalize hierarchically out of distribution (OOD). We study this shift using controlled grammar-learning tasks: question formation and tense inflection. We find a model learns to generalize hierarchically if its training data is complex-in particular, if it includes center-embedded clauses, a special syntactic structure. Under this definition, complex data drives hierarchical rules, while less complex data encourages shortcut learning in the form of n-gram-like linear rules. Furthermore, we find that a model uses rules to generalize, whether hierarchical or linear, if its training data is diverse-in particular, if it includes many distinct syntax trees in the training set. Under this definition, diverse data promotes stable rule learning, whereas less diverse data promotes memorization of individual syntactic sequences. Finally, intermediate diversity and intermediate complexity form an unstable regime, which is characterized by oscillatory learning dynamics and inconsistent behaviors across random seeds. These results highlight the central role of training data in shaping generalization and explain why competing strategies can lead to unstable outcomes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 001ac123-64df-4968-82c2-5be803350d4fCited by top-tier papers1
Ask how each one uses itBuilds on14
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language ModelsIsabel Papadimitriou, Dan JurafskyEMNLP 2020 · 40 citations
Related papers
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 17 citations
- Steering Language Models in Multi-Token Generation: A Case Study on Tense and AspectAlina Klerings, Jannik Brinkmann, Daniel Ruffinelli, Simone Paolo PonzettoEMNLP 2025
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita et al.EMNLP 2020 · 2 citations
- Learning curves theory for hierarchically compositional data with power-law distributed featuresFrancesco Cagnetta, Hyunmo Kang, Matthieu WyartICML 2025
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 4 citations
