Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
Siyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran Yang
Abstract
In-context learning (ICL) is a cornerstone of large language model (LLM) functionality, yet its theoretical foundations remain elusive due to the complexity of transformer architectures. In particular, most existing work only theoretically explains how the attention mechanism facilitates ICL under certain data models. It remains unclear how the other building blocks of the transformer contribute to ICL. To address this question, we study how a two-attention-layer transformer is trained to perform ICL on -gram Markov chain data, where each token in the Markov chain statistically depends on the previous tokens. We analyze a sophisticated transformer model featuring relative positional embedding, multi-head softmax attention, and a feed-forward layer with normalization. We prove that the gradient flow with respect to a cross-entropy ICL loss converges to a limiting model that performs a generalized version of the induction head mechanism with a learned feature, resulting from the congruous contribution of all the building blocks. In the limiting model, the first attention layer acts as a , copying past tokens within a given window to each position, and the feed-forward network with normalization acts as a that generates a feature vector by only looking at informationally relevant parents from the window. Finally, the second attention layer is a that compares these features with the feature at the output position, and uses the resulting similarity scores to generate the desired output. Our theory is further validated by experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02a28d05-4c01-41e2-bf41-c928b497d627Cited by top-tier papers35
- Enhancing Graph Transformers with Hierarchical Distance Structural EncodingYuankai Luo, Hongkang Li, Lei Shi, Xiao-Ming WuNeurIPS 2024 · 26 citations
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander et al.NeurIPS 2025 · 20 citations
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William GilpinICLR 2026 · 16 citations
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and ScaleJames A. Michaelov, Roger P. Levy, Benjamin BergenNeurIPS 2025 · 15 citations
- Symmetries in language statistics shape the geometry of model representationsDhruva Karkada, Daniel Korchinski, Andres Nava, Matthieu Wyart et al.ICML 2026 · 15 citations
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov ChainsChanakya Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman et al.NeurIPS 2025 · 4 citations
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 1 citation
- Selective induction Heads: How Transformers Select Causal Structures in ContextFrancesco D'Angelo, Francesco Croce, Nicolas FlammarionICLR 2025
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 117 citations
