Reinforcement Learning Teachers of Test Time Scaling
Edoardo Cetin, Tianyu Zhao, Yujin Tang
摘要
Training reasoning language models (LMs) with reinforcement learning (RL) for one-hot correctness inherently relies on the LM being able to explore and solve its task with some chance at initialization. Furthermore, a key use case of reasoning LMs is to act as teachers for distilling new students and cold-starting future RL iterations rather than being deployed themselves. From these considerations, we introduce a new framework that avoids RL's exploration challenge by training a new class of Reinforcement-Learned Teachers (RLTs) focused on yielding the most effective downstream distillation. RLTs are prompted with both the question and solution to each problem, and tasked to simply"connect-the-dots"with detailed explanations tailored for their students. We train RLTs with dense rewards obtained by feeding each explanation to the student and testing its understanding of the problem's solution. In practice, the raw outputs of a 7B RLT provide higher final performance on competition and graduate-level tasks than existing distillation and cold-starting pipelines that collect and postprocess the reasoning traces of orders of magnitude larger LMs. Furthermore, RLTs maintain their effectiveness when training larger students and when applied zero-shot to out-of-distribution tasks, unlocking new levels of efficiency and re-usability for the RL reasoning framework. Code available at: https://github.com/SakanaAI/RLT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun 等ICLR 2026 · 被引用 22 次
- Variational Reasoning for Language ModelsXiangxin Zhou, Zichen Liu, Haonan Wang, Chao Du 等ICLR 2026 · 被引用 6 次
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important ReasoningQinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin 等ICML 2026 · 被引用 4 次
- Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought GenerationGuangyue Peng, Zongchao Chen, Wen Luo, Yuntao Wen 等ICML 2026
它引用的顶会 Paper16
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory InversionZhen Tan, Chengshuai Zhao, Song Wang, Jundong Li 等ICLR 2026
- RLKD: Distilling LLMs' Reasoning via Reinforcement LearningShicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu 等AAAI 2026 · 被引用 2 次
- Think before Recommendation: Autonomous Reasoning-enhanced RecommenderXiaoyu Kong, Junguang Jiang, Bin Liu, Ziru Xu 等NeurIPS 2025 · 被引用 17 次
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language ModelsSiyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang 等ICML 2026 · 被引用 245 次
