How Transformers Learn Causal Structure with Gradient Descent
Eshaan Nichani, Alex Damian, Jason D. Lee
Abstract
The incredible success of transformers on sequence modeling tasks can be largely attributed to the self-attention mechanism, which allows information to be transferred between different parts of a sequence. Self-attention allows transformers to encode causal structure which makes them particularly suitable for sequence modeling. However, the process by which transformers learn such causal structure via gradient-based training algorithms remains poorly understood. To better understand this process, we introduce an in-context learning task that requires learning latent causal structure. We prove that gradient descent on a simplified two-layer transformer learns to solve this task by encoding the latent causal graph in the first attention layer. The key insight of our proof is that the gradient of the attention matrix encodes the mutual information between tokens. As a consequence of the data processing inequality, the largest entries of this gradient correspond to edges in the latent causal graph. As a special case, when the sequences are generated from in-context Markov chains, we prove that transformers learn an induction head (Olsson et al., 2022). We confirm our theoretical findings by showing that transformers trained on our in-context learning task are able to recover a wide variety of causal structures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d47fe69-feda-4ba4-a928-e08f4bae5032Cited by top-tier papers85
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasksTianyu He, Darshil Doshi, Aritra Das, Andrey GromovNeurIPS 2024 · 52 citations
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 48 citations
- In-Context Learning with Representations: Contextual Generalization of Trained TransformersTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2024 · 45 citations
- Iteration Head: A Mechanistic Study of Chain-of-ThoughtVivien Cabannes, Charles Arnal, Wassim Bouaziz, Xingyu Yang et al.NeurIPS 2024 · 44 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
Related papers
- Selective induction Heads: How Transformers Select Causal Structures in ContextFrancesco D'Angelo, Francesco Croce, Nicolas FlammarionICLR 2025
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov ChainsChanakya Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman et al.NeurIPS 2025 · 4 citations
- How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical GuaranteeJianzhe Wei, Siyu Chen, Jianliang He, Zhuoran YangICLR 2026
- Transformers Learn Latent Mixture Models In-Context via Mirror DescentFrancesco D'Angelo, Nicolas FlammarionICLR 2026 · 2 citations
