RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Kohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
Abstract
Large language models (LLMs) are typically trained by reinforcement learning (RL) with verifiable rewards (RLVR) and supervised fine-tuning (SFT) on reasoning traces to improve their reasoning abilities. However, how these methods shape reasoning capabilities remains largely elusive. Going beyond an accuracy-based investigation of how these two components sculpt the reasoning process, this paper introduces a novel analysis framework that quantifies reasoning paths and captures their qualitative changes under each training process (with models of 1.5B, 7B, and 14B parameters on mathematical and code domains). Specifically, we investigate the reasoning process at two levels of granularity: the trajectory-level, which examines complete reasoning outputs, and the step-level, which analyzes reasoning graphs whose nodes correspond to individual reasoning steps. Notably, clustering of unique reasoning trajectories shows complementary effects: RL compresses incorrect trajectories, whereas SFT expands correct ones. Step-level analysis reveals that RL steepens (about 2.5 times), while SFT flattens (reduced to about one-third), the decay rates of node visitation frequency, degree, and betweenness centrality distributions in the reasoning graph. This indicates that RL concentrates reasoning functionality into a small subset of steps, while SFT homogenizes it across many steps. Furthermore, by evaluating the reasoning graph topologies from multiple perspectives, we delineate the shared and distinct characteristics of RL and SFT. Our work presents a novel reasoning path perspective that explains why the current best practice of two-stage training, with SFT followed by RL, is successful, and offers practical implications for data construction and more efficient learning approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82b02c5f-4caa-4c35-beaf-ba5073ba824dCited by top-tier papers5
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao et al.ACL 2026 · 5 citations
- Optimizing Few-Step Generation with Adaptive Matching DistillationLichen Bai, zikai ZHOU, Shitong Shao, Wenliang Zhong et al.ICML 2026 · 4 citations
- RL makes MLLMs see better than SFTJunha Song, Sangdoo Yun, Dongyoon Han, Jaegul Choo et al.ICLR 2026 · 2 citations
- ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model MergingJunyao Yang, Chen Qian, Wen Shen, Yong Liu et al.ACL 2026 · 1 citation
- Modeling Hierarchical Thinking in Large Reasoning ModelsG M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-GhazalehICML 2026
Related papers
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model ReasoningLingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma et al.ICML 2026 · 1 citation
- Native Reasoning Models: Training Language Models to Reason on Unverifiable DataYuanfu Wang, Zhixuan Liu, Li xiangtian, Chaochao Lu et al.ICLR 2026 · 3 citations
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai et al.ICML 2026 · 24 citations
- InT: Self-Proposed Interventions Enable Credit Assignment in LLM ReasoningMatthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang et al.ICLR 2026 · 12 citations
