Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
Yihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang, Rujun Han, Gufeng Zhang, Yanfei Chen, Wei Wang, Tomas Pfister, Chen-Yu Lee
摘要
Large Language Models (LLMs) often struggle with problems that require multi-step reasoning. For small-scale open-source models, Reinforcement Learning with Verifiable Rewards (RLVR) fails when correct solutions are rarely sampled even after many attempts, while Supervised Fine-Tuning (SFT) tends to overfit long demonstrations through rigid token-by-token imitation. To address this gap, we propose Supervised Reinforcement Learning (SRL), a framework that reformulates problem solving as generating a sequence of logical "actions". SRL trains the model to generate an internal reasoning monologue before committing to each action. It provides smoother rewards based on the similarity between the model's actions and expert actions extracted from the SFT dataset in a step-wise manner. This supervision offers richer learning signals even when all rollouts are incorrect, while encouraging flexible reasoning guided by expert demonstrations. As a result, SRL enables small models to learn challenging problems previously unlearnable by SFT or RLVR. Moreover, initializing training with SRL before refining with RLVR yields the strongest overall performance. Beyond reasoning benchmarks, SRL generalizes effectively to agentic software engineering tasks, establishing it as a robust and versatile training framework for reasoning-oriented LLMs. Introduction Large Language Models (LLMs) have demonstrated remarkable versatility across diverse domains, ranging from mathematical problem-solving (Wang et al., 2025
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual GroundingHee Suk Yoon, Eunseop Yoon, Jaehyun Jang, SooHwan Eom 等ICML 2026 · 被引用 6 次
- Seg-ReSearch: Segmentation with Interleaved Reasoning and External SearchTianming Liang, Qirui Du, Jian-Fang Hu, Haichao Jiang 等ICML 2026 · 被引用 5 次
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsPeizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper15
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu 等ICML 2024 · 被引用 376 次
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
相关 Paper
- CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical ReasoningZijun Gao, Zhikun Xu, Xiao Ye, Ben ZhouICLR 2026
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution ReasoningLingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye 等ACL 2026 · 被引用 2 次
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai 等ICML 2026 · 被引用 24 次
- Incentivizing LLM Reasoning via Reinforcement Learning with Functional Monte Carlo Tree SearchKongcheng Zhang, QI YAO, Baisheng Lai, Jiaxing Huang 等ICLR 2026
