Context-level Language Modeling by Learning Predictive Context Embeddings
beiya dai, Yuliang Liu, Yunchong Song, Daozheng Xue, Qipeng Guo, Kai Chen, Xinbing Wang, Bowen Zhou, Zhouhan Lin
Abstract
We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible with standard autoregressive, token-by-token evaluation paradigms (e.g., perplexity). Extensive experiments with GPT-2 and Pythia backbones (up to 1.5B parameters and 300B training tokens) reveal that ContextLM shifts the Pareto frontier of scaling laws, exhibiting superior efficiency in parameters, training tokens, and FLOPs. Our results show that ContextLM could already achieve the baseline perplexity using 39% fewer parameters and demonstrates robust generalization improvements on extensive downstream tasks under equivalent parameter counts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen et al.ACL 2025 · 116 citations
Related papers
- Learned Meta-Tokens for Language ModelingAlok N. Shah, Khush Gupta, Keshav Ramji, Pratik ChaudhariICLR 2026 · 2 citations
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous SpaceBoyi Zeng, He Li, Shixiang Song, Yixuan Wang et al.ICML 2026 · 10 citations
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Predicting the Order of Upcoming Tokens Improves Language ModelingZayd Muhammad Kawakibi Zuhri, Erland Hilman Fuadi, Alham Fikri AjiICML 2026 · 3 citations
- Training Trajectories of Language Models Across ScalesMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin et al.ACL 2023 · 12 citations
