Programming over Thinking: Efficient and Robust Multi-Constraint Planning
Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang
Abstract
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints. Existing large language model (LLM) approaches face fundamental limitations in this domain. Pure reasoning paradigms, which rely on long natural language chains, are prone to inconsistency, error accumulation, and prohibitive cost as constraints compound. Conversely, LLMs combined with coding- or solver-based strategies lack flexibility: they often generate problem-specific code from scratch or depend on fixed solvers, failing to capture generalizable logic across diverse problems. To address these challenges, we introduce the Scalable COde Planning Engine (SCOPE), a framework that disentangles query-specific reasoning from generic code execution. By separating reasoning from execution, SCOPE produces solver functions that are consistent, deterministic, and reusable across queries while requiring only minimal changes to input parameters. SCOPE achieves state-of-the-art performance while lowering cost and latency. For example, with GPT-4o, it reaches 93.1% success on TravelPlanner, a 61.6% gain over the best baseline (CoT) while cutting inference cost by 1.4x and time by 4.67x. Code is available at https://github.com/DerrickGXD/SCOPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81d02d91-d3b1-4567-9b2b-8d79f882e015Related papers
- OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity ScalingYitian Chen, Cheng Cheng, Yinan Sun, Zi Ling et al.ICML 2026 · 3 citations
- CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form PlanningJiaxin Wen, Jian Guan, Hongning Wang, Wei Wu et al.ICLR 2025
- CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li et al.ACL 2024
- SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided SearchDong Li, Xujiang Zhao, Linlin Yu, Yanchi Liu et al.NeurIPS 2025 · 15 citations
- Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized ProgrammingYilun Hao, Yang Zhang, Chuchu FanICLR 2025
