Provable Long-Range Benefits of Next-Token Prediction
Xinyuan Cao, Santosh S. Vempala
摘要
Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token prediction is provably powerful for learning longer-range structure, even with commonly used neural network architectures. Specifically, we prove that optimizing next-token prediction over a Recurrent Neural Network yields a model that closely approximates the training distribution: for held-out documents sampled from the training distribution, no algorithm of bounded description length limited to examining the next k tokens, for any k, can distinguish between k consecutive tokens of such documents and k tokens generated by the learned language model following the same prefix. We provide polynomial bounds (in k, independent of the document length) on the model size needed to achieve such k-token indistinguishability, offering a complexity-theoretic explanation for the long-range coherence observed in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
相关 Paper
- Auto-Regressive Next-Token Predictors are Universal LearnersEran MalachICML 2024 · 被引用 65 次
- Coherence boosting: When your pretrained language model is not paying enough attentionNikolay Malkin, Zhen Wang, Nebojsa JojicACL 2022 · 被引用 45 次
- Fractal Patterns May Illuminate the Success of Next-Token PredictionIbrahim M. Alabdulmohsin, Vinh Q. Tran, Mostafa DehghaniNeurIPS 2024 · 被引用 8 次
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade 等ICML 2025
- Do Long-Range Language Models Actually Use Long-Range Context?Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit IyyerEMNLP 2021 · 被引用 35 次
