ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning
Lingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye, Zhongxin Liu, Xiaoxue Ren, Lingfeng Bao
Abstract
Code LLMs still struggle with code execution reasoning, especially in smaller models. Existing methods rely on supervised fine-tuning (SFT) with teacher-generated explanations, primarily in two forms: (1) input-output (I/O) prediction chains and (2) natural-language descriptions of execution traces. However, intermediate execution steps cannot be explicitly verified during SFT, so the training objective can be reduced to merely matching teacher explanations. Moreover, training data is typically collected without explicit control over task difficulty. We introduce ExecVerify, which goes beyond text imitation by incorporating verifiable white-box rewards derived from execution traces, including next-statement prediction and variable value/type prediction. Our work first builds a dataset with multiple difficulty levels via constraint-based program synthesis. Then, we apply reinforcement learning (RL) to reward correct answers about both intermediate execution steps and final outputs, aligning the training objective with semantic correctness at each execution step. Finally, we adopt a twostage training pipeline that first enhances execution reasoning and then transfers to code generation. Experiments demonstrate that a 7B model trained with ExecVerify achieves performance comparable to 32B models on code reasoning benchmarks and improves pass@1 by up to 5.9% on code generation tasks over strong post-training baselines 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1512232d-9374-4489-ba20-84c7339a4f5dCited by top-tier papers1
Ask how each one uses itBuilds on9
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding et al.ICML 2024 · 246 citations
- Neural Program Repair with Execution-based BackpropagationHe Ye, Matias Martinez, Martin MonperrusICSE 2022 · 146 citations
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik et al.ICLR 2024 · 45 citations
Related papers
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang et al.ICLR 2026 · 11 citations
- LeDex: Training LLMs to Better Self-Debug and Explain CodeNan Jiang, Xiaopeng Li, Shiqi Wang, Qiang Zhou et al.NeurIPS 2024 · 6 citations
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
- Reasoning Through Execution: Unifying Process and Outcome Rewards for Code GenerationZhuohao Yu, Weizheng Gu, Yidong Wang, Xingru Jiang et al.ICML 2025
- CODERL+: Improving Code Generation via Reinforcement with Execution Semantics AlignmentXue Jiang, Yihong Dong, Mengyang Liu, Hongyi Deng et al.ACL 2026 · 18 citations
