Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
Yitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge, Yinyu Ye
摘要
Optimization modeling is fundamental to decision-making in fields such as supply chain management, logistics, and financial engineering, but its complexity presents a major barrier to adoption. Automating model creation from natural language is key to improving efficiency and access. However, while Large Language Models (LLMs) are a promising tool for this, they often produce flawed or infeasible results due to errors and hallucinations. To address this issue, we propose Solver-Informed Reinforcement Learning (SIRL), a framework that uses Reinforcement Learning with Verifiable Reward to improve LLMs ability to generate accurate and executable optimization models. Specifically, SIRL automatically assesses the executable code and the instance-level mathematical model represented by the associated .lp files. This process yields precise feedback on syntactic validity, feasibility, and solution quality, which serves as a direct reward signal to guide the reinforcement learning process. Furthermore, this verification mechanism also supports our instance-enhanced self-consistency method for creating high-quality training data. Extensive experiments on diverse public benchmarks demonstrate that models trained with our SIRL framework achieve state-of-the-art performance, substantially outperforming existing methods in generating accurate and executable optimization models. Specifically, our SIRL-32B model surpasses DeepSeek-V3 and OpenAI-o3 on the majority of these benchmarks. Our code is publicly available at https://github.com/Cardinal-Operations/SIRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language ModelsChenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong GeICLR 2026 · 被引用 29 次
- BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge EvaluationPeng Lai, Zhihao Ou, Yong Wang, Longyue Wang 等ICLR 2026 · 被引用 15 次
- Constructing Industrial-Scale Optimization Modeling BenchmarkZhong Li, Hongliang Lu, Tao Wei, Yuxuan Chen 等ICML 2026 · 被引用 5 次
- Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side VerificationHaoyang Liu, Jie Wang, Boxuan Niu, Xiongwei Han 等ICML 2026 · 被引用 4 次
- OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity ScalingYitian Chen, Cheng Cheng, Yinan Sun, Zi Ling 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- Large Language Models as End-to-end Combinatorial Optimization SolversXia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao 等NeurIPS 2025 · 被引用 37 次
- DeepOR: A Deep Reasoning Foundation Model for Optimization ModelingZiyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan 等AAAI 2026 · 被引用 1 次
- VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based VerificationSumaya Abdul Rahman, Seckhen Cuellar, Ghani Raissov, Mohammad RazaICML 2026 · 被引用 1 次
- SAC-Opt: Semantic Anchors for Iterative Correction in Optimization ModelingYansen Zhang, Qingcan Kang, Yujie chen, Yufei Wang 等ICML 2026
- MURKA: Multi-Reward Reinforcement Learning with Knowledge Alignment for Optimization TasksWantong Xie, Yi-Xiang Hu, Jieyang Xu, Feng Wu 等NeurIPS 2025 · 被引用 8 次
