Transformers on Markov data: Constant depth suffices
Nived Rajaraman, Marco Bondaschi, Ashok Vardhan Makkuva, Kannan Ramchandran, Michael Gastpar
摘要
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transformers on data drawn from Markov processes, where the conditional distribution of the next symbol in a sequence depends on the previous symbols observed. We observe a surprising phenomenon empirically which contradicts previous findings: when trained for sufficiently long, a transformer with a fixed depth and head per layer is able to achieve low test loss on sequences drawn from Markov sources, even as grows. Furthermore, this low test loss is achieved by the transformer's ability to represent and learn the in-context conditional empirical distribution. On the theoretical side, our main result is that a transformer with a single head and three layers can represent the in-context conditional empirical distribution for Markov sources, concurring with our empirical observations. Along the way, we prove that attention-only transformers with layers can represent the in-context conditional empirical distribution by composing induction heads to track the previous symbols in the sequence. These results provide more insight into our current understanding of the mechanisms by which transformers learn to capture context, by understanding their behavior on Markov sources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach 等NeurIPS 2024 · 被引用 140 次
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 被引用 48 次
- Local to Global: Learning Dynamics and Effect of Initialization for TransformersAshok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle 等NeurIPS 2024 · 被引用 16 次
- From Markov to Laplace: How Mamba In-Context Learns Markov ChainsMarco Bondaschi, Nived Rajaraman, Xiuying Wei, Razvan Pascanu 等ICLR 2026 · 被引用 10 次
- Latent Concept Disentanglement in Transformer-based Language ModelsGuanzhe Hong, Bhavya Vasudeva, Vatsal Sharan, Cyrus Rashtchian 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
相关 Paper
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov ChainsChanakya Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman 等NeurIPS 2025 · 被引用 4 次
- Attention with Markov: A Curious Case of Single-layer TransformersAshok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle 等ICLR 2025
- An Analysis of Tokenization: Transformers under Markov DataNived Rajaraman, Jiantao Jiao, Kannan RamchandranNeurIPS 2024 · 被引用 16 次
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 被引用 117 次
- What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning TasksXingwu Chen, Difan ZouICML 2024 · 被引用 22 次
