The Pitfalls of Next-Token Prediction
Gregor Bachmann, Vaishnavh Nagarajan
Abstract
Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective. As a starting point, we argue that the two often-conflated phases of next-token predictionautoregressive inference and teacher-forced training -must be treated distinctly. The popular criticism that errors can compound during autoregressive inference, crucially assumes that teacherforcing has learned an accurate next-token predictor. This assumption sidesteps a more deep-rooted problem we expose: in certain classes of tasks, teacher-forcing can simply fail to learn an accurate next-token predictor in the first place. We describe a general mechanism of how teacherforcing can fail, and design a minimal planning task where both the Transformer and the Mamba architecture empirically fail in that manner -remarkably, despite the task being straightforward to learn. Finally, we provide preliminary evidence that this failure can be resolved using teacherless training, a simple modification using dummy tokens that predicts multiple tokens in advance. We hope this finding can ground future debates and inspire explorations beyond the next-token prediction paradigm. We make our code available under https://github.com/gregorbachmann/ Next-Token-Failures * Equal contribution 1 ETH Zürich, Switzerland 2 Google Research, US.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adf59fef-1e2e-4cff-b320-c2ce34e34cc3Cited by top-tier papers82
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-FoldAmrith Setlur, Saurabh Garg, Xinyang Geng, Naman Garg et al.NeurIPS 2024 · 143 citations
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
- Transformers Represent Belief State Geometry in their Residual StreamAdam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen et al.NeurIPS 2024 · 83 citations
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesJin Wang, Yao Lai, Aoxue Li, Shifeng Zhang et al.NeurIPS 2025 · 45 citations
Builds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Predicting the Order of Upcoming Tokens Improves Language ModelingZayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri AjiICML 2026 · 3 citations
- ALPINE: Unveiling The Planning Capability of Autoregressive Learning in Language ModelsSiwei Wang, Yifei Shen, Shi Feng, Haoran Sun et al.NeurIPS 2024 · 18 citations
- Semformer: Transformer Language Models with Semantic PlanningYongjing Yin, Junran Ding, Kai Song, Yue ZhangEMNLP 2024
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- TokenUnify: Scaling Up Autoregressive Pretraining for Neuron SegmentationYinda Chen, Haoyuan Shi, Xiaoyu Liu, Te Shi et al.ICCV 2025 · 1 citation
