Provable Long-Range Benefits of Next-Token Prediction
Xinyuan Cao, Santosh S. Vempala
Abstract
Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token prediction is provably powerful for learning longer-range structure, even with commonly used neural network architectures. Specifically, we prove that optimizing next-token prediction over a Recurrent Neural Network yields a model that closely approximates the training distribution: for held-out documents sampled from the training distribution, no algorithm of bounded description length limited to examining the next k tokens, for any k, can distinguish between k consecutive tokens of such documents and k tokens generated by the learned language model following the same prefix. We provide polynomial bounds (in k, independent of the document length) on the model size needed to achieve such k-token indistinguishability, offering a complexity-theoretic explanation for the long-range coherence observed in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4913af0f-7bb4-4d2f-8115-84b2dcf75c66Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
Related papers
- Auto-Regressive Next-Token Predictors are Universal LearnersEran MalachICML 2024 · 65 citations
- Coherence boosting: When your pretrained language model is not paying enough attentionNikolay Malkin, Zhen Wang, Nebojsa JojicACL 2022 · 45 citations
- Fractal Patterns May Illuminate the Success of Next-Token PredictionIbrahim M. Alabdulmohsin, Vinh Q. Tran, Mostafa DehghaniNeurIPS 2024 · 8 citations
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade et al.ICML 2025
- Do Long-Range Language Models Actually Use Long-Range Context?Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit IyyerEMNLP 2021 · 35 citations
