StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
Chenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong Ge
Abstract
Large Language Models (LLMs) have shown promising capabilities for solving Operations Research (OR) problems. While reinforcement learning serves as a powerful paradigm for LLM training on OR problems, existing works generally face two key limitations. First, outcome reward suffers from the , where correct final answers can reinforce flawed reasoning. Second, conventional discriminative process supervision is , failing to evaluate the interdependent steps of OR modeling holistically. To this end, we introduce \textbf{\texttt{StepORLM}}, a novel self-evolving framework with generative process supervision. At its core, features a co-evolutionary loop where a policy model and a generative process reward model (GenPRM) iteratively improve on each other. This loop is driven by a dual-feedback mechanism: definitive, outcome-based verification from an external solver, and nuanced, holistic process evaluation from the GenPRM. The combined signal is used to align the policy via Weighted Direct Preference Optimization (W-DPO) and simultaneously refine the GenPRM. Our resulting 8B-parameter establishes a new state-of-the-art across six benchmarks, significantly outperforming vastly larger generalist models, agentic methods, and specialized baselines. Moreover, the co-evolved GenPRM is able to act as a powerful and universally applicable process verifier, substantially boosting the inference scaling performance of both our own model and other existing LLMs. We release our models and code to facilitate future research (https://github.com/0xzhouchenyu/StepORLM).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb2d2ade-0ee5-41cc-975a-5c035515abd4Cited by top-tier papers4
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function CallingJianghao Lin, Yuanyuan Shi, Xin Peng, Renjie Ding et al.ACL 2026 · 3 citations
- Strategy-Aware Optimization Modeling with Reasoning LLMsRuiqing Zhao, Fengzhi Li, Yuan Zuo, Rui Liu et al.ICML 2026 · 1 citation
- A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and UsageCongmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen et al.ACL 2026
- FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear ProgrammingHongpei Li, Hui Yuan, Han Zhang, Jianghao Lin et al.ICLR 2026
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- ReEvo: Large Language Models as Hyper-Heuristics with Reflective EvolutionHaoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto et al.NeurIPS 2024 · 424 citations
- Chain-of-Experts: When LLMs Meet Complex Operations Research ProblemsZiyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu et al.ICLR 2024 · 136 citations
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsAli AhmadiTeshnizi, Wenzhi Gao, Madeleine UdellICML 2024 · 77 citations
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization ModelingYitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge et al.NeurIPS 2025 · 54 citations
Related papers
- Reasoning Through Execution: Unifying Process and Outcome Rewards for Code GenerationZhuohao Yu, Weizheng Gu, Yidong Wang, Xingru Jiang et al.ICML 2025
- OR-PRM: A Process Reward Model for Algorithmic Problem in Operations ResearchYilin Wang, Heng Zhou, Dongxing Mao, Linjie Li et al.ICLR 2026
- RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningKaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong et al.NeurIPS 2025 · 44 citations
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou et al.AAAI 2026 · 68 citations
- DeepOR: A Deep Reasoning Foundation Model for Optimization ModelingZiyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan et al.AAAI 2026 · 1 citation
