StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
Chenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong Ge
摘要
Large Language Models (LLMs) have shown promising capabilities for solving Operations Research (OR) problems. While reinforcement learning serves as a powerful paradigm for LLM training on OR problems, existing works generally face two key limitations. First, outcome reward suffers from the , where correct final answers can reinforce flawed reasoning. Second, conventional discriminative process supervision is , failing to evaluate the interdependent steps of OR modeling holistically. To this end, we introduce \textbf{\texttt{StepORLM}}, a novel self-evolving framework with generative process supervision. At its core, features a co-evolutionary loop where a policy model and a generative process reward model (GenPRM) iteratively improve on each other. This loop is driven by a dual-feedback mechanism: definitive, outcome-based verification from an external solver, and nuanced, holistic process evaluation from the GenPRM. The combined signal is used to align the policy via Weighted Direct Preference Optimization (W-DPO) and simultaneously refine the GenPRM. Our resulting 8B-parameter establishes a new state-of-the-art across six benchmarks, significantly outperforming vastly larger generalist models, agentic methods, and specialized baselines. Moreover, the co-evolved GenPRM is able to act as a powerful and universally applicable process verifier, substantially boosting the inference scaling performance of both our own model and other existing LLMs. We release our models and code to facilitate future research (https://github.com/0xzhouchenyu/StepORLM).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function CallingJianghao Lin, Yuanyuan Shi, Xin Peng, Renjie Ding 等ACL 2026 · 被引用 3 次
- Strategy-Aware Optimization Modeling with Reasoning LLMsRuiqing Zhao, Fengzhi Li, Yuan Zuo, Rui Liu 等ICML 2026 · 被引用 1 次
- A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and UsageCongmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen 等ACL 2026
- FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear ProgrammingHongpei Li, Hui Yuan, Han Zhang, Jianghao Lin 等ICLR 2026
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- ReEvo: Large Language Models as Hyper-Heuristics with Reflective EvolutionHaoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto 等NeurIPS 2024 · 被引用 424 次
- Chain-of-Experts: When LLMs Meet Complex Operations Research ProblemsZiyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu 等ICLR 2024 · 被引用 136 次
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsAli AhmadiTeshnizi, Wenzhi Gao, Madeleine UdellICML 2024 · 被引用 77 次
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization ModelingYitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge 等NeurIPS 2025 · 被引用 54 次
相关 Paper
- Reasoning Through Execution: Unifying Process and Outcome Rewards for Code GenerationZhuohao Yu, Weizheng Gu, Yidong Wang, Xingru Jiang 等ICML 2025
- OR-PRM: A Process Reward Model for Algorithmic Problem in Operations ResearchYilin Wang, Heng Zhou, Dongxing Mao, Linjie Li 等ICLR 2026
- RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningKaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong 等NeurIPS 2025 · 被引用 44 次
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou 等AAAI 2026 · 被引用 68 次
- DeepOR: A Deep Reasoning Foundation Model for Optimization ModelingZiyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan 等AAAI 2026 · 被引用 1 次
