Towards Understanding the Universality of Transformers for Next-Token Prediction
Michael Eli Sander, Gabriel Peyré
Abstract
Causal Transformers are trained to predict the next token for a given context. While it is widely accepted that self-attention is crucial for encoding the causal structure of sequences, the precise underlying mechanism behind this in-context autoregressive learning ability remains unclear. In this paper, we take a step towards understanding this phenomenon by studying the approximation ability of Transformers for nexttoken prediction. Specifically, we explore the capacity of causal Transformers to predict the next token x t+1 given an autoregressive sequence (x 1 , . . . , x t ) as a prompt, where x t+1 = f (x t ), and f is a context-dependent function that varies with each sequence. On the theoretical side, we focus on specific instances, namely when f is linear or when (x t ) t≥1 is periodic. We explicitly construct a Transformer (with linear, exponential, or softmax attention) that learns the mapping f in-context through a causal kernel descent method. The causal kernel descent method we propose provably estimates x t+1 based solely on past and current observations (x 1 , . . . , x t ), with connections to the Kaczmarz algorithm in Hilbert spaces. We present experimental results that validate our theoretical findings and suggest their applicability to more general mappings f .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d63a561e-2090-45f1-9eb0-2ea6146659efCited by top-tier papers3
- Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for ReasoningXinhao Yao, Ruifeng Ren, Yun Liao, Lizhong Ding et al.ICLR 2026 · 6 citations
- A Theoretical Analysis of Detecting Large Model-Generated Time SeriesJunji Hou, Junzhou Zhao, Shuo Zhang, Pinghui WangAAAI 2026 · 2 citations
- On learning linear dynamical systems in context with attention layersMaria-Luiza Vladarean, Xuhui Zhang, Suvrit SraICLR 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
Related papers
- On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and CapabilityChenyu Zheng, Wei Huang, Rongzhen Wang, Guoqiang Wu et al.NeurIPS 2024 · 10 citations
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 117 citations
- How do Transformers Perform In-Context Autoregressive Learning ?Michael Eli Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel et al.ICML 2024 · 21 citations
- Transformers are Universal In-context LearnersTakashi Furuya, Maarten V. de Hoop, Gabriel PeyréICLR 2025 · 2 citations
- One-Layer Transformer Provably Learns One-Nearest Neighbor In ContextZihao Li, Yuan Cao, Cheng Gao, Yihan He et al.NeurIPS 2024 · 25 citations
