Pre-Training Curriculum for Multi-Token Prediction in Language Models
Ansar Aynetdinov, Alan Akbik
Abstract
Multi-token prediction (MTP) is a recently proposed pre-training objective for language models. Rather than predicting only the next token (NTP), MTP predicts the next k tokens at each prediction step, using multiple prediction heads. MTP has shown promise in improving downstream performance, inference speed, and training efficiency, particularly for large models. However, prior work has shown that smaller language models (SLMs) struggle with the MTP objective. To address this, we propose a curriculum learning strategy for MTP training, exploring two variants: a forward curriculum, which gradually increases the complexity of the pre-training objective from NTP to MTP, and a reverse curriculum, which does the opposite. Our experiments show that the forward curriculum enables SLMs to better leverage the MTP objective during pre-training, improving downstream NTP performance and generative output quality, while retaining the benefits of self-speculative decoding. The reverse curriculum achieves stronger NTP performance and output quality, but fails to provide any selfspeculative decoding benefits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang et al.ACL 2020 · 156 citations
- LayerSkip: Enabling Early Exit Inference and Self-Speculative DecodingMostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer et al.ACL 2024 · 22 citations
Related papers
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- Predicting the Order of Upcoming Tokens Improves Language ModelingZayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri AjiICML 2026 · 3 citations
- Parallel Token Prediction for Language ModelsFelix Draxler, Justus C. Will, Farrin Marouf Sofian, Theofanis Karaletsos et al.ICLR 2026 · 6 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesDivyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki et al.ICLR 2026 · 15 citations
