RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Kohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
摘要
Large language models (LLMs) are typically trained by reinforcement learning (RL) with verifiable rewards (RLVR) and supervised fine-tuning (SFT) on reasoning traces to improve their reasoning abilities. However, how these methods shape reasoning capabilities remains largely elusive. Going beyond an accuracy-based investigation of how these two components sculpt the reasoning process, this paper introduces a novel analysis framework that quantifies reasoning paths and captures their qualitative changes under each training process (with models of 1.5B, 7B, and 14B parameters on mathematical and code domains). Specifically, we investigate the reasoning process at two levels of granularity: the trajectory-level, which examines complete reasoning outputs, and the step-level, which analyzes reasoning graphs whose nodes correspond to individual reasoning steps. Notably, clustering of unique reasoning trajectories shows complementary effects: RL compresses incorrect trajectories, whereas SFT expands correct ones. Step-level analysis reveals that RL steepens (about 2.5 times), while SFT flattens (reduced to about one-third), the decay rates of node visitation frequency, degree, and betweenness centrality distributions in the reasoning graph. This indicates that RL concentrates reasoning functionality into a small subset of steps, while SFT homogenizes it across many steps. Furthermore, by evaluating the reasoning graph topologies from multiple perspectives, we delineate the shared and distinct characteristics of RL and SFT. Our work presents a novel reasoning path perspective that explains why the current best practice of two-stage training, with SFT followed by RL, is successful, and offers practical implications for data construction and more efficient learning approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao 等ACL 2026 · 被引用 5 次
- Optimizing Few-Step Generation with Adaptive Matching DistillationLichen Bai, zikai ZHOU, Shitong Shao, Wenliang Zhong 等ICML 2026 · 被引用 4 次
- RL makes MLLMs see better than SFTJunha Song, Sangdoo Yun, Dongyoon Han, Jaegul Choo 等ICLR 2026 · 被引用 2 次
- ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model MergingJunyao Yang, Chen Qian, Wen Shen, Yong Liu 等ACL 2026 · 被引用 1 次
- Modeling Hierarchical Thinking in Large Reasoning ModelsG M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-GhazalehICML 2026
相关 Paper
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model ReasoningLingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma 等ICML 2026 · 被引用 1 次
- Native Reasoning Models: Training Language Models to Reason on Unverifiable DataYuanfu Wang, Zhixuan Liu, Li xiangtian, Chaochao Lu 等ICLR 2026 · 被引用 3 次
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai 等ICML 2026 · 被引用 24 次
- InT: Self-Proposed Interventions Enable Credit Assignment in LLM ReasoningMatthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang 等ICLR 2026 · 被引用 12 次
