Attention with Markov: A Curious Case of Single-layer Transformers
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, Michael Gastpar
Abstract
Attention-based transformers have achieved tremendous success across a variety of disciplines including natural languages. To deepen our understanding of their sequential modeling capabilities, there is a growing interest in using Markov input processes to study them. A key finding is that when trained on first-order Markov chains, transformers with two or more layers consistently develop an induction head mechanism to estimate the in-context bigram conditional distribution. In contrast, single-layer transformers, unable to form an induction head, directly learn the Markov kernel but often face a surprising challenge: they become trapped in local minima representing the unigram distribution, whereas deeper models reliably converge to the ground-truth bigram. While single-layer transformers can theoretically model first-order Markov chains, their empirical failure to learn this simple kernel in practice remains a curious phenomenon. To explain this contrasting behavior of single-layer models, in this paper we introduce a new framework for a principled analysis of transformers via Markov chains. Leveraging our framework, we theoretically characterize the loss landscape of single-layer transformers and show the existence of global minima (bigram) and bad local minima (unigram) contingent on data properties and model architecture. We precisely delineate the regimes under which these local optima occur. Backed by experiments, we demonstrate that our theoretical findings are in congruence with the empirical results. Finally, we outline several open problems in this arena. Code is available at https://github.com/Bond1995/Markov.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7eacfdc-ee02-442b-9ffe-3646f435b4a9Cited by top-tier papers12
- From Markov to Laplace: How Mamba In-Context Learns Markov ChainsMarco Bondaschi, Nived Rajaraman, Xiuying Wei, Razvan Pascanu et al.ICLR 2026 · 10 citations
- Nonparametric Teaching of Attention LearnersChen Zhang, Jianghui Wang, Bingyang Cheng, Zhongtao Chen et al.ICLR 2026 · 3 citations
- Decoupling Positional and Symbolic Attention in TransformersFelipe Urrutia, Jorge Salas, Alexander Kozachinskiy, Cristian Buc Calderon et al.ICLR 2026 · 3 citations
- Multiple Choice Learning of Low-Rank Adapters for Language ModelingVictor Letzelter, Hugo Malard, Mathieu Fontaine, Gaël Richard et al.ICML 2026 · 1 citation
- Towards Understanding Transformers in Learning Random WalksWei Shi, Yuan CaoNeurIPS 2025 · 1 citation
Builds on29
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
Related papers
- Local to Global: Learning Dynamics and Effect of Initialization for TransformersAshok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle et al.NeurIPS 2024 · 16 citations
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov ChainsChanakya Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman et al.NeurIPS 2025 · 4 citations
- Transformers on Markov data: Constant depth sufficesNived Rajaraman, Marco Bondaschi, Ashok Vardhan Makkuva, Kannan Ramchandran et al.NeurIPS 2024 · 33 citations
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- An Analysis of Tokenization: Transformers under Markov DataNived Rajaraman, Jiantao Jiao, Kannan RamchandranNeurIPS 2024 · 16 citations
