Program Synthesis Guided Reinforcement Learning for Partially Observed Environments
Yichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu, Armando Solar-Lezama, Martin C. Rinard
Abstract
A key challenge for reinforcement learning is solving long-horizon planning problems. Recent work has leveraged programs to guide reinforcement learning in these settings. However, these approaches impose a high manual burden on the user since they must provide a guiding program for every new task. Partially observed environments further complicate the programming task because the program must implement a strategy that correctly, and ideally optimally, handles every possible configuration of the hidden regions of the environment. We propose a new approach, model predictive program synthesis (MPPS), that uses program synthesis to automatically generate the guiding programs. It trains a generative model to predict the unobserved portions of the world, and then synthesizes a program based on samples from this model in a way that is robust to its uncertainty. In our experiments, we show that our approach significantly outperforms non-program-guided approaches on a set of challenging benchmarks, including a 2D Minecraft-inspired environment where the agent must complete a complex sequence of subtasks to achieve its goal, and achieves a similar performance as using handcrafted programs to guide the agent. Our results demonstrate that our approach can obtain the benefits of program-guided reinforcement learning without requiring the user to provide a new guiding program for every new task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6519dfc-0269-4f7f-a5fd-de89400caf35Cited by top-tier papers13
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 59 citations
- GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisYushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang et al.NeurIPS 2022 · 23 citations
- Automated Program Refinement: Guide and Verify Code Large Language Model with Refinement CalculusYufan Cai, Zhe Hou, David Sanán, Xiaokun Luan et al.POPL 2025 · 20 citations
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi et al.NeurIPS 2023 · 15 citations
- SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning SystemsOubo Ma, Yuwen Pu, Linkang Du, Yang Dai et al.CCS 2024 · 6 citations
Builds on16
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 126 citations
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 112 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learningKevin Ellis, Catherine Wong, Maxwell I. Nye, Mathias Sablé-Meyer et al.PLDI 2021 · 97 citations
- Neurosymbolic Reinforcement Learning with Formally Verified ExplorationGreg Anderson, Abhinav Verma, Isil Dillig, Swarat ChaudhuriNeurIPS 2020 · 91 citations
Related papers
- E-MAPP: Efficient Multi-Agent Reinforcement Learning with Parallel Program GuidanceCan Chang, Ni Mu, Jiajun Wu, Ling Pan et al.NeurIPS 2022 · 14 citations
- HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisShraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick et al.NeurIPS 2024 · 34 citations
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller et al.NeurIPS 2025 · 31 citations
- Reinforcement Learning and Data-Generation for Syntax-Guided SynthesisJulian Parsert, Elizabeth PolgreenAAAI 2024 · 7 citations
- ProTo: Program-Guided Transformer for Program-Guided TasksZelin Zhao, Karan Samel, Binghong Chen, Le SongNeurIPS 2021 · 37 citations
