Lune

STOC2026Top-tier venue

Provable Long-Range Benefits of Next-Token Prediction

Xinyuan Cao, Santosh S. Vempala

2026Year

Abstract

Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token prediction is provably powerful for learning longer-range structure, even with commonly used neural network architectures. Specifically, we prove that optimizing next-token prediction over a Recurrent Neural Network yields a model that closely approximates the training distribution: for held-out documents sampled from the training distribution, no algorithm of bounded description length limited to examining the next k tokens, for any k, can distinguish between k consecutive tokens of such documents and k tokens generated by the learned language model following the same prefix. We provide polynomial bounds (in k, independent of the document length) on the model size needed to achieve such k-token indistinguishability, offering a complexity-theoretic explanation for the long-range coherence observed in practice.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4913af0f-7bb4-4d2f-8115-84b2dcf75c66

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines