Planning in Natural Language Improves LLM Search for Code Generation
Evan Z. Wang, Federico Cassano, Catherine Wu, Yunfeng Bai, William Song, Vaskar Nath, Ziwen Han, Sean M. Hendryx, Summer Yue, Hugh Zhang
摘要
While scaling training compute has led to remarkable improvements in large language models (LLMs), scaling inference compute has not yet yielded analogous gains. We hypothesize that a core missing component is a lack of diverse LLM outputs, leading to inefficient search due to models repeatedly sampling highly similar, yet incorrect generations. We empirically demonstrate that this lack of diversity can be mitigated by searching over candidate plans for solving a problem in natural language. Based on this insight, we propose PLANSEARCH, a novel search algorithm which shows strong results across HumanEval+, MBPP+, and LiveCodeBench (a contamination-free benchmark for competitive coding). PLANSEARCH generates a diverse set of observations about the problem and uses these observations to construct plans for solving the problem. By searching over plans in natural language rather than directly over code solutions, PLANSEARCH explores a significantly more diverse range of potential solutions compared to baseline search methods. Using PLANSEARCH on top of Claude 3.5 Sonnet achieves a pass@200 of 77.0% on LiveCodeBench, outperforming both the best pass-rate achieved without any search (pass@1 = 41.4%) and using standard repeated sampling on top of existing non-search models (pass@200 = 60.6%). Finally, we show that, across all models, search algorithms, and benchmarks analyzed, we can accurately predict performance gains from search as a function of the diversity over generated ideas. Code can be found at https://github.com/scaleapi/plansearch .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- CodeDPO: Aligning Code Models with Self Generated and Verified Source CodeKechi Zhang, Ge Li, Yihong Dong, Jingjing Xu 等ACL 2025 · 被引用 45 次
- Searching Latent Program SpacesMatthew Macfarlane, Clément BonnetNeurIPS 2025 · 被引用 23 次
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksVishnu Sarukkai, Zhiqiang Xie, Kayvon FatahalianNeurIPS 2025 · 被引用 22 次
- RPG: A Repository Planning Graph for Unified and Scalable Codebase GenerationJane Luo, Xin Zhang, Steven Liu, Jie Wu 等ICLR 2026 · 被引用 18 次
- Generalizable Heuristic Generation Through LLMs with Meta-OptimizationYiding Shi, Jianan Zhou, Wen Song, Jieyi Bi 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
相关 Paper
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li 等ICLR 2025
- Thought of Search: Planning with Language Models Through The Lens of EfficiencyMichael Katz, Harsha Kokel, Kavitha Srinivas, Shirin SohrabiNeurIPS 2024 · 被引用 50 次
- DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree TraversalVaibhav Aggarwal, Ojasv Kamal, Abhinav Japesh, Zhijing Jin 等ACL 2025
- Planning with Large Language Models for Code GenerationShun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding 等ICLR 2023 · 被引用 15 次
- DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction TuningYejie Wang, Keqing He, Guanting Dong, Pei Wang 等ACL 2024
