Execution-guided within-prompt search for programming-by-example
Gust Verbruggen, Ashish Tiwari, Mukul Singh, Vu Le, Sumit Gulwani
Abstract
Large language models (LLMs) can generate code from examples without being limited to a DSL, but they lack search, as sampled programs are independent. In this paper, we use an LLM as a policy that generates lines of code and then join these lines of code to let the LLM implicitly estimate the value of each of these lines in its next iteration. We further guide the policy and value estimation by executing each line and annotating it with its results on the given examples. This allows us to search for programs within a single (expanding) prompt until a sound program is found, by letting the policy reason in both the syntactic (code) and semantic (execution) space. We evaluate within-prompt search on straight-line Python code generation using five benchmarks across different domains (strings, lists, and arbitrary Python programming problems). We show that the model uses the execution results to guide the search and that within-prompt search performs well at low token budgets. We also analyze how the model behaves as a policy and value, show that it can parallelize the search, and that it can implicitly backtrack over earlier generations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bcbe5b7-e28f-4766-b1cb-d5f39c83a9acCited by top-tier papers3
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim et al.NeurIPS 2025 · 4 citations
- PRM-PBE: Process Reward Model for Reinforcement Learning in Programming-by-ExampleYue Fang, Zhi Jin, Jie An, Hongshen Chen et al.ICML 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu et al.ICLR 2024 · 156 citations
Related papers
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
- Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python CodeAugusto B. Corrêa, André Grahl Pereira, Jendrik SeippNeurIPS 2025 · 27 citations
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code GenerationHenrijs Princis, Arindam Sharma, Cristina DavidPLDI 2026
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang et al.FSE 2024 · 68 citations
- What Makes Large Language Models Reason in (Multi-Turn) Code Generation?Kunhao Zheng, Juliette Decugis, Jonas Gehring, Taco Cohen et al.ICLR 2025
