Planning in Natural Language Improves LLM Search for Code Generation
Evan Z. Wang, Federico Cassano, Catherine Wu, Yunfeng Bai, William Song, Vaskar Nath, Ziwen Han, Sean M. Hendryx, Summer Yue, Hugh Zhang
Abstract
While scaling training compute has led to remarkable improvements in large language models (LLMs), scaling inference compute has not yet yielded analogous gains. We hypothesize that a core missing component is a lack of diverse LLM outputs, leading to inefficient search due to models repeatedly sampling highly similar, yet incorrect generations. We empirically demonstrate that this lack of diversity can be mitigated by searching over candidate plans for solving a problem in natural language. Based on this insight, we propose PLANSEARCH, a novel search algorithm which shows strong results across HumanEval+, MBPP+, and LiveCodeBench (a contamination-free benchmark for competitive coding). PLANSEARCH generates a diverse set of observations about the problem and uses these observations to construct plans for solving the problem. By searching over plans in natural language rather than directly over code solutions, PLANSEARCH explores a significantly more diverse range of potential solutions compared to baseline search methods. Using PLANSEARCH on top of Claude 3.5 Sonnet achieves a pass@200 of 77.0% on LiveCodeBench, outperforming both the best pass-rate achieved without any search (pass@1 = 41.4%) and using standard repeated sampling on top of existing non-search models (pass@200 = 60.6%). Finally, we show that, across all models, search algorithms, and benchmarks analyzed, we can accurately predict performance gains from search as a function of the diversity over generated ideas. Code can be found at https://github.com/scaleapi/plansearch .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46367c86-a162-4a26-8340-5e6d2ed1ee13Cited by top-tier papers33
- CodeDPO: Aligning Code Models with Self Generated and Verified Source CodeKechi Zhang, Ge Li, Yihong Dong, Jingjing Xu et al.ACL 2025 · 45 citations
- Searching Latent Program SpacesMatthew Macfarlane, Clément BonnetNeurIPS 2025 · 23 citations
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksVishnu Sarukkai, Zhiqiang Xie, Kayvon FatahalianNeurIPS 2025 · 22 citations
- RPG: A Repository Planning Graph for Unified and Scalable Codebase GenerationJane Luo, Xin Zhang, Steven Liu, Jie Wu et al.ICLR 2026 · 18 citations
- Generalizable Heuristic Generation Through LLMs with Meta-OptimizationYiding Shi, Jianan Zhou, Wen Song, Jieyi Bi et al.ICLR 2026 · 14 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
Related papers
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li et al.ICLR 2025
- Thought of Search: Planning with Language Models Through The Lens of EfficiencyMichael Katz, Harsha Kokel, Kavitha Srinivas, Shirin SohrabiNeurIPS 2024 · 50 citations
- DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree TraversalVaibhav Aggarwal, Ojasv Kamal, Abhinav Japesh, Zhijing Jin et al.ACL 2025
- Planning with Large Language Models for Code GenerationShun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding et al.ICLR 2023 · 15 citations
- DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction TuningYejie Wang, Keqing He, Guanting Dong, Pei Wang et al.ACL 2024
