RLP: Reinforcement as a Pretraining Objective
Ali Hatamizadeh, Syeda Nahida Akter, Shrimai Prabhumoye, Jan Kautz, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi
Abstract
The dominant paradigm for training large reasoning models starts with pre-training using next-token prediction loss on vast amounts of data. Reinforcement learning, while powerful in scaling reasoning, is introduced only as the very last phase of post-training, preceded by supervised fine-tuning. While dominant, is this an optimal way of training? In this paper, we present RLP, an information-driven reinforcement pretraining objective, that brings the core spirit of reinforcement learning---exploration---to the last phase of pretraining. The key idea is to treat chain-of-thought as an exploratory action, with rewards computed based on the information gain it provides for predicting future tokens. This training objective essentially encourages the model to think for itself before predicting what comes next, thus teaching an independent thinking behavior earlier in the pretraining. More concretely, the reward signal measures the increase in log-likelihood of the next token when conditioning on both context and a sampled reasoning chain, compared to conditioning on context alone. This approach yields a verifier-free dense reward signal, allowing for efficient training for the full document stream during pretraining. Specifically, RLP reframes reinforcement learning for reasoning as a pretraining objective on ordinary text, bridging the gap between next-token prediction and the emergence of useful chain-of-thought reasoning. Pretraining with RLP on Qwen3-1.7B-Base lifts the overall average across an eight‑benchmark math‑and‑science suite by 19%. With identical post‑training, the gains compound, with the largest improvements on reasoning‑heavy tasks such as AIME25 and MMLU‑Pro. Applying RLP to the hybrid NVIDIA-Nemotron-Nano-12B-v2-Base increases the overall average from 42.81% to 61.32% and raises the average on scientific reasoning by 23%, demonstrating scalability across architectures and model sizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ca383ba-be35-4cf3-9f9a-6b23d79013daCited by top-tier papers4
- Efficient Process Reward Modeling via Contrastive Mutual InformationNakyung Lee, Sangwoo Hong, Jungwoo LeeACL 2026 · 4 citations
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann et al.ICML 2026
- PretrainZero: Reinforcement Active Learning on Pretraining DataXingrun Xing, Zhiyuan Fan, Jie Lou, Guoqi Li et al.ICML 2026
- InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model PretrainingShirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad et al.ICML 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
Related papers
- Reinforcement Learning on Pre-Training DataSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang et al.ACL 2026 · 11 citations
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL SynergyZihan Liu, Zhuolin Yang, Yang Chen, Chankyu Lee et al.ICLR 2026 · 73 citations
- Reinforcement Mid-TrainingYijun Tian, Shaoyu Chen, Zhichao Xu, Yawei Wang et al.ICLR 2026 · 4 citations
- Native Reasoning Models: Training Language Models to Reason on Unverifiable DataYuanfu Wang, Zhixuan Liu, Li xiangtian, Chaochao Lu et al.ICLR 2026 · 3 citations
- MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningZhiheng Xi, Yuhui Wang, Yiwen Ding, Guanyu Li et al.AAAI 2026
