Reinforcement Learning Teachers of Test Time Scaling
Edoardo Cetin, Tianyu Zhao, Yujin Tang
Abstract
Training reasoning language models (LMs) with reinforcement learning (RL) for one-hot correctness inherently relies on the LM being able to explore and solve its task with some chance at initialization. Furthermore, a key use case of reasoning LMs is to act as teachers for distilling new students and cold-starting future RL iterations rather than being deployed themselves. From these considerations, we introduce a new framework that avoids RL's exploration challenge by training a new class of Reinforcement-Learned Teachers (RLTs) focused on yielding the most effective downstream distillation. RLTs are prompted with both the question and solution to each problem, and tasked to simply"connect-the-dots"with detailed explanations tailored for their students. We train RLTs with dense rewards obtained by feeding each explanation to the student and testing its understanding of the problem's solution. In practice, the raw outputs of a 7B RLT provide higher final performance on competition and graduate-level tasks than existing distillation and cold-starting pipelines that collect and postprocess the reasoning traces of orders of magnitude larger LMs. Furthermore, RLTs maintain their effectiveness when training larger students and when applied zero-shot to out-of-distribution tasks, unlocking new levels of efficiency and re-usability for the RL reasoning framework. Code available at: https://github.com/SakanaAI/RLT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d74cac5-7fe4-4c07-9df5-c41adbd7abbeCited by top-tier papers4
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun et al.ICLR 2026 · 22 citations
- Variational Reasoning for Language ModelsXiangxin Zhou, Zichen Liu, Haonan Wang, Chao Du et al.ICLR 2026 · 6 citations
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important ReasoningQinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin et al.ICML 2026 · 4 citations
- Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought GenerationGuangyue Peng, Zongchao Chen, Wen Luo, Yuntao Wen et al.ICML 2026
Builds on16
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory InversionZhen Tan, Chengshuai Zhao, Song Wang, Jundong Li et al.ICLR 2026
- RLKD: Distilling LLMs' Reasoning via Reinforcement LearningShicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu et al.AAAI 2026 · 2 citations
- Think before Recommendation: Autonomous Reasoning-enhanced RecommenderXiaoyu Kong, Junguang Jiang, Bin Liu, Ziru Xu et al.NeurIPS 2025 · 17 citations
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language ModelsSiyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang et al.ICML 2026 · 245 citations
