Lune

ICLR2025Top-tier venue

Towards Understanding the Universality of Transformers for Next-Token Prediction

Michael Eli Sander, Gabriel Peyré

2025Year
3Top-tier citations

Abstract

Causal Transformers are trained to predict the next token for a given context. While it is widely accepted that self-attention is crucial for encoding the causal structure of sequences, the precise underlying mechanism behind this in-context autoregressive learning ability remains unclear. In this paper, we take a step towards understanding this phenomenon by studying the approximation ability of Transformers for nexttoken prediction. Specifically, we explore the capacity of causal Transformers to predict the next token x t+1 given an autoregressive sequence (x 1 , . . . , x t ) as a prompt, where x t+1 = f (x t ), and f is a context-dependent function that varies with each sequence. On the theoretical side, we focus on specific instances, namely when f is linear or when (x t ) t≥1 is periodic. We explicitly construct a Transformer (with linear, exponential, or softmax attention) that learns the mapping f in-context through a causal kernel descent method. The causal kernel descent method we propose provably estimates x t+1 based solely on past and current observations (x 1 , . . . , x t ), with connections to the Kaczmarz algorithm in Hilbert spaces. We present experimental results that validate our theoretical findings and suggest their applicability to more general mappings f .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d63a561e-2090-45f1-9eb0-2ea6146659ef

Cited by top-tier papers3

Ask how each one uses it

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines