Better & Faster Large Language Models via Multi-token Prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
摘要
Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at once results in higher sample efficiency. More specifically, at each position in the training corpus, we ask the model to predict the following n tokens using n independent output heads, operating on top of a shared model trunk. Considering multi-token prediction as an auxiliary training task, we measure improved downstream capabilities with no overhead in training time for both code and natural language models. The method is increasingly useful for larger model sizes, and keeps its appeal when training for multiple epochs. Gains are especially pronounced on generative benchmarks like coding, where our models consistently outperform strong baselines by several percentage points. Our 13B parameter models solves 12 % more problems on HumanEval and 17 % more on MBPP than comparable next-token models. Experiments on small algorithmic tasks demonstrate that multi-token prediction is favorable for the development of induction heads and algorithmic reasoning capabilities. As an additional benefit, models trained with 4-token prediction are up to 3 times faster at inference, even with large batch sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper94
- MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsSijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng 等NeurIPS 2024 · 被引用 125 次
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 等NeurIPS 2025 · 被引用 90 次
- Transformers Represent Belief State Geometry in their Residual StreamAdam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen 等NeurIPS 2024 · 被引用 83 次
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at ScaleChunrui Han, Guopeng Li, Jingwei Wu, Quan Sun 等ICLR 2026 · 被引用 58 次
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 被引用 48 次
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- HyperTree Proof Search for Neural Theorem ProvingGuillaume Lample, Timothée Lacroix, Marie-Anne Lachaux, Aurélien Rodriguez 等NeurIPS 2022 · 被引用 271 次
相关 Paper
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang 等NeurIPS 2025 · 被引用 16 次
- Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesDivyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki 等ICLR 2026 · 被引用 15 次
- Context-level Language Modeling by Learning Predictive Context Embeddingsbeiya dai, Yuliang Liu, Yunchong Song, Daozheng Xue 等ICML 2026 · 被引用 5 次
- Predicting the Order of Upcoming Tokens Improves Language ModelingZayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri AjiICML 2026 · 被引用 3 次
- Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingRaghavv Goel, Mukul Gagrani, Mingu Lee, Christopher LottICML 2026
