OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling
Yitian Chen, Cheng Cheng, Yinan Sun, Zi Ling, Dongdong Ge
Abstract
We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reasoning and precise formulation. To this end, we introduce OPT-ENGINE, an extensible benchmark framework with quantifiable and controllable complexity. OPT-ENGINE spans ten canonical Operations Research problems, systematically scaling from Linear Programming to Mixed-Integer Programming, providing a structured environment to probe the limits of automated problem formulation and solving. Utilizing OPT-Engine, we address three pivotal research questions. First, we examine whether Pure-Text Reasoning (PTR) via classical Chain-of-Thought can efficiently tackle optimization tasks, finding that PTR suffers from a critical robustness gap as task complexity increases. Second, we examine whether integrating external computational tools can mitigate PTR's arithmetic weaknesses and improve performance. Our results indicate that while such tools help with local calculations, they still fail to adhere to global optimization constraints. Finally, we pinpoint that for the current SOTA paradigm, Solver-integrated Reasoning (SIR), the automated formulation of constraints represents the primary bottleneck. These findings clarify the limitations of current paradigms and provide a structured roadmap for developing next-generation LLMs for optimization modeling. We release our code and data to facilitate future research ( https://github.com/Cardinal- Operations/OPTEngine).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton et al.NeurIPS 2025 · 507 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization ModelingYitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge et al.NeurIPS 2025 · 54 citations
Related papers
- Evaluating LLM Reasoning in the Operations Research Domain with ORQAMahdi Mostajabdaveh, Timothy Tin Long Yu, Samarendra Chandan Bindu Dash, Rindra Ramamonjison et al.AAAI 2025 · 1 citation
- Chain-of-Experts: When LLMs Meet Complex Operations Research ProblemsZiyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu et al.ICLR 2024 · 136 citations
- SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided SearchDong Li, Xujiang Zhao, Linlin Yu, Yanchi Liu et al.NeurIPS 2025 · 15 citations
- Programming over Thinking: Efficient and Robust Multi-Constraint PlanningDerrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen et al.ACL 2026
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsAli AhmadiTeshnizi, Wenzhi Gao, Madeleine UdellICML 2024 · 77 citations
