Semformer: Transformer Language Models with Semantic Planning
Yongjing Yin, Junran Ding, Kai Song, Yue Zhang
摘要
Next-token prediction serves as the dominant component in current neural language models. During the training phase, the model employs teacher forcing, which predicts tokens based on all preceding ground truth tokens. However, this approach has been found to create shortcuts, utilizing the revealed prefix to spuriously fit future tokens, potentially compromising the accuracy of the next-token predictor. In this paper, we introduce Semformer, a novel method of training a Transformer language model that explicitly models the semantic planning of response. Specifically, we incorporate a sequence of planning tokens into the prefix, guiding the planning token representations to predict the latent semantic representations of the response, which are induced by an autoencoder. In a minimal planning task (i.e., graph path-finding), our model exhibits near-perfect performance and effectively mitigates shortcut learning, a feat that standard training methods and baseline models have been unable to accomplish. Furthermore, we pretrain Semformer from scratch with 125M parameters, demonstrating its efficacy through measures of perplexity, in-context learning, and fine-tuning on summarization tasks 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesDivyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki 等ICLR 2026 · 被引用 15 次
- Multi-Token Prediction Needs RegistersAnastasios Gerontopoulos, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 被引用 13 次
- Toward Consistent World Models with Multi-Token Prediction and Latent Semantic EnhancementQimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou 等ACL 2026 · 被引用 3 次
- Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More MoreArvid FrydenlundACL 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
相关 Paper
- Predicting the Order of Upcoming Tokens Improves Language ModelingZayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri AjiICML 2026 · 被引用 3 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Planning-Aware Code Infilling via Horizon-Length PredictionYifeng Ding, Hantian Ding, Shiqi Wang, Qing Sun 等EMNLP 2025
- NITP: Next Implicit Token Prediction for LLM Pre-trainingXiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin 等ICML 2026
- Next-ToBE: Probabilistic Next Token-Bag Exploitation for Activating Anticipatory Capacity in LLMsYihe Liu, Huibin Wang, Xianming Hu, Pinyi Zhang 等ICLR 2026
