Think before you speak: Training Language Models With Pause Tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, Vaishnavh Nagarajan
摘要
Language models generate responses by producing a series of tokens in immediate succession: the token is an outcome of manipulating hidden vectors per layer, one vector per preceding token. What if instead we were to let the model manipulate say, hidden vectors, before it outputs the token? We operationalize this idea by performing training and inference on language models with a (learnable) token, a sequence of which is appended to the input prefix. We then delay extracting the model's outputs until the last pause token is seen, thereby allowing the model to process extra computation before committing to an answer. We empirically evaluate on decoder-only models of 1B and 130M parameters with causal pretraining on C4, and on downstream tasks covering reasoning, question-answering, general understanding and fact recall. Our main finding is that inference-time delays show gains when the model is both pre-trained and finetuned with delays. For the 1B model, we witness gains on 8 of 9 tasks, most prominently, a gain of EM score on the QA task of SQuAD, on CommonSenseQA and accuracy on the reasoning task of GSM8k. Our work raises a range of conceptual and practical future research questions on making delayed next-token prediction a widely applicable new paradigm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper99
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton 等NeurIPS 2025 · 被引用 507 次
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim 等NeurIPS 2025 · 被引用 143 次
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen 等CVPR 2026 · 被引用 124 次
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
相关 Paper
- Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn MoreXialie Zhuang, Zhikai Jia, Jianjin Li, Zhenyu Zhang 等ICML 2025
- Optimal Splitting of Language Models from Mixtures to Specialized DomainsSkyler Seto, Pierre Ablin, Anastasiia Filippova, Jiayuan Ye 等ICML 2026 · 被引用 2 次
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceHoujun Liu, Shikhar Murty, Christopher Manning, Róbert CsordásICML 2026 · 被引用 3 次
- Deliberation in Latent Space via Differentiable Cache AugmentationLuyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie 等ICML 2025
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck 等ACL 2022
