Program Synthesis Guided Reinforcement Learning for Partially Observed Environments
Yichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu, Armando Solar-Lezama, Martin C. Rinard
摘要
A key challenge for reinforcement learning is solving long-horizon planning problems. Recent work has leveraged programs to guide reinforcement learning in these settings. However, these approaches impose a high manual burden on the user since they must provide a guiding program for every new task. Partially observed environments further complicate the programming task because the program must implement a strategy that correctly, and ideally optimally, handles every possible configuration of the hidden regions of the environment. We propose a new approach, model predictive program synthesis (MPPS), that uses program synthesis to automatically generate the guiding programs. It trains a generative model to predict the unobserved portions of the world, and then synthesizes a program based on samples from this model in a way that is robust to its uncertainty. In our experiments, we show that our approach significantly outperforms non-program-guided approaches on a set of challenging benchmarks, including a 2D Minecraft-inspired environment where the agent must complete a complex sequence of subtasks to achieve its goal, and achieves a similar performance as using handcrafted programs to guide the agent. Our results demonstrate that our approach can obtain the benefits of program-guided reinforcement learning without requiring the user to provide a new guiding program for every new task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 被引用 59 次
- GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisYushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang 等NeurIPS 2022 · 被引用 23 次
- Automated Program Refinement: Guide and Verify Code Large Language Model with Refinement CalculusYufan Cai, Zhe Hou, David Sanán, Xiaokun Luan 等POPL 2025 · 被引用 20 次
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi 等NeurIPS 2023 · 被引用 15 次
- SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning SystemsOubo Ma, Yuwen Pu, Linkang Du, Yang Dai 等CCS 2024 · 被引用 6 次
它引用的顶会 Paper16
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 被引用 126 次
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 被引用 112 次
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 被引用 97 次
- DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learningKevin Ellis, Catherine Wong, Maxwell I. Nye, Mathias Sablé-Meyer 等PLDI 2021 · 被引用 97 次
- Neurosymbolic Reinforcement Learning with Formally Verified ExplorationGreg Anderson, Abhinav Verma, Isil Dillig, Swarat ChaudhuriNeurIPS 2020 · 被引用 91 次
相关 Paper
- E-MAPP: Efficient Multi-Agent Reinforcement Learning with Parallel Program GuidanceCan Chang, Ni Mu, Jiajun Wu, Ling Pan 等NeurIPS 2022 · 被引用 14 次
- HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisShraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick 等NeurIPS 2024 · 被引用 34 次
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller 等NeurIPS 2025 · 被引用 31 次
- Reinforcement Learning and Data-Generation for Syntax-Guided SynthesisJulian Parsert, Elizabeth PolgreenAAAI 2024 · 被引用 7 次
- ProTo: Program-Guided Transformer for Program-Guided TasksZelin Zhao, Karan Samel, Binghong Chen, Le SongNeurIPS 2021 · 被引用 37 次
