Reinforcement Mid-Training
Yijun Tian, Shaoyu Chen, Zhichao Xu, Yawei Wang, Jinhe Bi, Peng Han, Wei Wang
Abstract
The development of state-of-the-art large language models is commonly understood as a two-stage process involving pre-training and post-training. We point out the need for an additional intermediate stage called reinforcement mid-training with potential for strong performance gains. In this paper, we formally define the problem and identify three key challenges: (1) inefficient training due to excessive reasoning steps, (2) disregard of the imbalanced token entropy distribution, and (3) underutilization of token information. To address these challenges, we propose RMT, a framework for efficient, adaptive, and unified reinforcement mid-training with various innovative components. In particular, we first introduce a dynamic token budget mechanism that constrains unnecessary reasoning steps and mitigates model overthinking. Next, we design a curriculum-based adaptive sampling method that fosters a progressive learning trajectory from easy to hard tokens. Finally, we present a dual training strategy that combines reinforcement learning with next-token prediction, ensuring targeted learning on key tokens and full exploitation of all token information. Extensive experiments demonstrate the superiority of RMT over state-of-the-art methods, achieving up to +64.91% performance improvement with only 21% of the reasoning length in language modeling. We also show that checkpoints obtained after reinforcement mid-training can benefit the subsequent post-training, yielding up to +18.76% improvement in the mathematical domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- EchoRL: Reinforcement Learning via Rollout EchoingJinhe Bi, Aniri -, Minglai Yang, Xingcheng Zhou et al.ICML 2026 · 6 citations
- Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language ModelsShaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen et al.ACL 2026 · 1 citation
- The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory EvolutionJinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang et al.ICML 2026
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- RLP: Reinforcement as a Pretraining ObjectiveAli Hatamizadeh, Syeda Nahida Akter, Shrimai Prabhumoye, Jan Kautz et al.ICLR 2026 · 26 citations
- UFT: Unifying Supervised and Reinforcement Fine-TuningMingyang Liu, Gabriele Farina, Asuman OzdaglarNeurIPS 2025 · 61 citations
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 58 citations
- Incentivizing LLM Reasoning via Reinforcement Learning with Functional Monte Carlo Tree SearchKongcheng Zhang, QI YAO, Baisheng Lai, Jiaxing Huang et al.ICLR 2026
- Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsMengqi Liao, Xiangyu Xi, Ruinian Chen, Jia Leng et al.EMNLP 2025 · 16 citations
