Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin, Christopher Potts
Abstract
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model posttraining recipes. A common assumption is that the reasoning chains trained through RLVR reliably represent how a model gets to its answer. In this paper, we develop two metrics for critically examining this assumption: Causal Importance of Reasoning (CIR), which measures the cumulative effect of reasoning tokens on the final answer, and Sufficiency of Reasoning (SR), which measures whether a verifier can arrive at an unambiguous answer based on the reasoning alone. Through experiments with the Qwen2.5 model series and ReasoningGym tasks, we find that: (1) While RLVR does improve task accuracy, it does not reliably improve CIR or SR, calling the role of reasoning in model performance into question. (2) A small amount of SFT before RLVR can be a remedy for low CIR and SR. (3) CIR and SR can be improved even without SFT by applying auxiliary CIR/SR rewards on top of the outcome-based reward. This joint reward matches the accuracy of RLVR while also leading to causally important and sufficient reasoning. These results show that RLVR does not always lead models to rely on reasoning in the way that is commonly thought, but this issue can be remedied with simple modifications to the post-training procedure. Our code is available at https: //github.com/yuqinan/cir-sr/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f84ea9b8-de67-44fb-8ac1-a8fc53e1fc68Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye et al.ICLR 2026 · 279 citations
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao et al.ICLR 2026 · 28 citations
Related papers
- Generalization of RLVR Using Causal Reasoning as a TestbedZhichu Lu, Hongyu Zhao, Shuo Sun, Hao Peng et al.ICLR 2026 · 4 citations
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use InsteadFeiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica et al.ICLR 2026 · 27 citations
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended GenerationYuxin Jiang, Yufei Wang, Qiyuan Zhang, Xingshan Zeng et al.ICLR 2026 · 5 citations
- Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningXiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li et al.NeurIPS 2025 · 19 citations
