The Grammar-Learning Trajectories of Neural Language Models
Leshem Choshen, Guy Hacohen, Daphna Weinshall, Omri Abend
摘要
The learning trajectories of linguistic phenomena in humans provide insight into linguistic representation, beyond what can be gleaned from inspecting the behavior of an adult speaker. To apply a similar approach to analyze neural language models (NLM), it is first necessary to establish that different models are similar enough in the generalizations they make. In this paper, we show that NLMs with different initialization, architecture, and training data acquire linguistic phenomena in a similar order, despite their different end performance. These findings suggest that there is some mutual inductive bias that underlies these models’ learning of linguistic phenomena. Taking inspiration from psycholinguistics, we argue that studying this inductive bias is an opportunity to study the linguistic representation implicit in NLMs.Leveraging these findings, we compare the relative performance on different phenomena at varying learning stages with simpler reference models. Results suggest that NLMs exhibit consistent “developmental” stages. Moreover, we find the learning trajectory to be approximately one-dimensional: given an NLM with a certain overall performance, it is possible to predict what linguistic generalizations it has already acquired.Initial analysis of these stages presents phenomena clusters (notably morphological ones), whose performance progresses in unison, suggesting a potential link between the generalizations behind them.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 被引用 304 次
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 被引用 163 次
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt 等ICLR 2024 · 被引用 119 次
- Not All Tokens Are What You Need for PretrainingZhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 等NeurIPS 2024 · 被引用 99 次
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
它引用的顶会 Paper7
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 被引用 163 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- A Theoretical Analysis of the Repetition Problem in Text GenerationZihao Fu, Wai Lam, Anthony Man-Cho So, Bei ShiAAAI 2021 · 被引用 114 次
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 被引用 87 次
- What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional EncodingYu-An Wang, Yun-Nung ChenEMNLP 2020 · 被引用 71 次
相关 Paper
- Examining the Inductive Bias of Neural Language Models with Artificial LanguagesJennifer C. White, Ryan CotterellACL 2021
- Emergence of a High-Dimensional Abstraction Phase in Language TransformersEmily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco 等ICLR 2025 · 被引用 1 次
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda 等ACL 2024
- Language Model Evaluation Beyond PerplexityClara Meister, Ryan CotterellACL 2021
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 被引用 9 次
