Data Drives Unstable Hierarchical Generalization in LMs
Tian Qin, Naomi Saphra, David Alvarez-Melis
摘要
Early in training, LMs can behave like n-gram models, but eventually, they often learn treebased syntactic rules and generalize hierarchically out of distribution (OOD). We study this shift using controlled grammar-learning tasks: question formation and tense inflection. We find a model learns to generalize hierarchically if its training data is complex-in particular, if it includes center-embedded clauses, a special syntactic structure. Under this definition, complex data drives hierarchical rules, while less complex data encourages shortcut learning in the form of n-gram-like linear rules. Furthermore, we find that a model uses rules to generalize, whether hierarchical or linear, if its training data is diverse-in particular, if it includes many distinct syntax trees in the training set. Under this definition, diverse data promotes stable rule learning, whereas less diverse data promotes memorization of individual syntactic sequences. Finally, intermediate diversity and intermediate complexity form an unstable regime, which is characterized by oscillatory learning dynamics and inconsistent behaviors across random seeds. These results highlight the central role of training data in shaping generalization and explain why competing strategies can lead to unstable outcomes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 被引用 186 次
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt 等ICLR 2024 · 被引用 119 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language ModelsIsabel Papadimitriou, Dan JurafskyEMNLP 2020 · 被引用 40 次
相关 Paper
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 被引用 17 次
- Steering Language Models in Multi-Token Generation: A Case Study on Tense and AspectAlina Klerings, Jannik Brinkmann, Daniel Ruffinelli, Simone Paolo PonzettoEMNLP 2025
- Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsEthan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita 等EMNLP 2020 · 被引用 2 次
- Learning curves theory for hierarchically compositional data with power-law distributed featuresFrancesco Cagnetta, Hyunmo Kang, Matthieu WyartICML 2025
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 被引用 4 次
