OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling
Zhicheng Yang, Yiwei Wang, Yinya Huang, Zhijiang Guo, Wei Shi, Xiongwei Han, Liang Feng, Linqi Song, Xiaodan Liang, Jing Tang
摘要
Large language models (LLMs) have exhibited their problem-solving abilities in mathematical reasoning. Solving realistic optimization (OPT) problems in application scenarios requires advanced and applied mathematics ability. However, current OPT benchmarks that merely solve linear programming are far from complex realistic situations. In this work, we propose OPTIBENCH, a benchmark for end-to-end optimization problem-solving with human-readable inputs and outputs. OPTIBENCH contains rich optimization problems, including linear and nonlinear programming with or without tabular data, which can comprehensively evaluate LLMs' solving ability. In our benchmark, LLMs are required to call a code solver to provide precise numerical answers. Furthermore, to alleviate the data scarcity for optimization problems, and to bridge the gap between open-source LLMs on a small scale (e.g., Llama-3-8b) and closed-source LLMs (e.g., GPT-4), we further propose a data synthesis method namely ReSocratic. Unlike general data synthesis methods that proceed from questions to answers, ReSocratic first incrementally synthesizes formatted optimization demonstrations with mathematical formulations step by step and then back-translates the generated demonstrations into questions. Based on this, we synthesize the RESOCRATIC-29K dataset. We further conduct supervised fine-tuning with RESOCRATIC-29K on multiple open-source models. Experimental results show that RESOCRATIC-29K significantly improves the performance of open-source models. The code and data can be found at https://github.com/yangzhch6/ReSocratic .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization ModelingYitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge 等NeurIPS 2025 · 被引用 54 次
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language ModelsChenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong GeICLR 2026 · 被引用 29 次
- MM-Agent: LLM as Agents for Real-world Mathematical Modeling ProblemFan Liu, Zherui Yang, Cancheng Liu, Tianrui Song 等NeurIPS 2025 · 被引用 28 次
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial OptimizationWeiwei Sun, Shengyu Feng, Shanda Li, Yiming YangAAAI 2026 · 被引用 20 次
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience LibraryMinwei Kong, Ao Qu, Xiaotong Guo, Wenbin Ouyang 等KDD 2026 · 被引用 16 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu 等ICLR 2024 · 被引用 558 次
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski 等ICLR 2023 · 被引用 171 次
- Chain-of-Experts: When LLMs Meet Complex Operations Research ProblemsZiyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu 等ICLR 2024 · 被引用 136 次
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsAli AhmadiTeshnizi, Wenzhi Gao, Madeleine UdellICML 2024 · 被引用 77 次
相关 Paper
- OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization ModelingHongliang Lu, Zhonglin Xie, Yaoyu Wu, Can Ren 等ICML 2025
- Evaluating LLM Reasoning in the Operations Research Domain with ORQAMahdi Mostajabdaveh, Timothy Tin Long Yu, Samarendra Chandan Bindu Dash, Rindra Ramamonjison 等AAAI 2025 · 被引用 1 次
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
- OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity ScalingYitian Chen, Cheng Cheng, Yinan Sun, Zi Ling 等ICML 2026 · 被引用 3 次
