Predicting the Emergence of Induction Heads in Language Model Pretraining
Tatsuya Aoyama, Ethan G Wilcox, Nathan Schneider
摘要
Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in the context of language modeling, remains wanting. In this study, we investigate the relationship between statistical properties of the training data and IH formation in both natural and synthetic training data settings. We show that: (1) a simple equation combining batch size and context size predicts the point at which IHs form and that this emergence point is agnostic to model size; (2) surface bigram repetition frequency and reliability strongly affect the formation of IHs, and we find an effective decision boundary in terms of these two values; (3) local dependency with high bigram repetition frequency and reliability is sufficient for IH formation, but categoriality and the shape of the marginal distribution appear to modulate IH formation near the decision boundary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach 等NeurIPS 2024 · 被引用 140 次
相关 Paper
- The mechanistic basis of data dependence and abrupt learning in an in-context classification taskGautam ReddyICLR 2024 · 被引用 112 次
- KV Shifting Attention Enhances Language ModelingMingyu Xu, Bingning Wang, Weipeng ChenICML 2025
- What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formationAaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan 等ICML 2024 · 被引用 77 次
- The emergence of sparse attention: impact of data distribution and benefits of repetitionNicolas Zucchet, Francesco D'Angelo, Andrew Kyle Lampinen, Stephanie ChanNeurIPS 2025 · 被引用 28 次
- Language Models Grow Less Humanlike beyond Phase TransitionTatsuya Aoyama, Ethan WilcoxACL 2025
