Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
Max Liu, Chan-Hung Yu, Wei-Hsu Lee, Cheng-Wei Hung, Yen-Chun Chen, Shao-Hua Sun
Abstract
Programmatic reinforcement learning (PRL) has been explored for representing policies through programs as a means to achieve interpretability and generalization. Despite promising outcomes, current state-of-the-art PRL methods are hindered by sample inefficiency, necessitating tens of millions of program-environment interactions. To tackle this challenge, we introduce a novel LLM-guided search framework (LLM-GS). Our key insight is to leverage the programming expertise and common sense reasoning of LLMs to enhance the efficiency of assumption-free, random-guessing search methods. We address the challenge of LLMs' inability to generate precise and grammatically correct programs in domain-specific languages (DSLs) by proposing a Pythonic-DSL strategy -an LLM is instructed to initially generate Python codes and then convert them into DSL programs. To further optimize the LLM-generated programs, we develop a search algorithm named Scheduled Hill Climbing, designed to efficiently explore the programmatic search space to improve the programs consistently. Experimental results in the Karel domain demonstrate our LLM-GS framework's superior effectiveness and efficiency. Extensive ablation studies further verify the critical role of our Pythonic-DSL strategy and Scheduled Hill Climbing algorithm. Moreover, we conduct experiments with two novel tasks, showing that LLM-GS enables users without programming skills and knowledge of the domain or DSL to describe the tasks in natural language to obtain performant programs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9cbe3cd-f863-4ca9-82c5-5f5c5fec1b91Cited by top-tier papers5
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token LearningChao-Chung Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Vivian Chen et al.NeurIPS 2025 · 19 citations
- Partition to Evolve: Niching-enhanced Evolution with LLMs for Automated Algorithm DiscoveryQinglong Hu, Qingfu ZhangNeurIPS 2025 · 13 citations
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control PoliciesQinglong Hu, Tong Xialiang, Mingxuan Yuan, Fei Liu et al.ICLR 2026 · 7 citations
- Hierarchical Programmatic Option FrameworkYu-An Lin, Chen-Tao Lee, Chih-Han Yang, Guan-Ting Liu et al.NeurIPS 2024 · 7 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
- Hierarchical Programmatic Reinforcement Learning via Learning to Compose ProgramsGuan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-Yi Lee et al.ICML 2023 · 21 citations
- HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisShraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick et al.NeurIPS 2024 · 34 citations
- Execution-guided within-prompt search for programming-by-exampleGust Verbruggen, Ashish Tiwari, Mukul Singh, Vu Le et al.ICLR 2025
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 6 citations
