Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
Tales Henrique Carvalho, Kenneth Tjhia, Levi Lelis
摘要
Recent works have introduced LEAPS and HPRL, systems that learn latent spaces of domain-specific languages, which are used to define programmatic policies for partially observable Markov decision processes (POMDPs). These systems induce a latent space while optimizing losses such as the behavior loss, which aim to achieve locality in program behavior, meaning that vectors close in the latent space should correspond to similarly behaving programs. In this paper, we show that the programmatic space, induced by the domain-specific language and requiring no training, presents values for the behavior loss similar to those observed in latent spaces presented in previous work. Moreover, algorithms searching in the programmatic space significantly outperform those in LEAPS and HPRL. To explain our results, we measured the "friendliness" of the two spaces to local search algorithms. We discovered that algorithms are more likely to stop at local maxima when searching in the latent space than when searching in the programmatic space. This implies that the optimization topology of the programmatic space, induced by the reward function in conjunction with the neighborhood function, is more conducive to search than that of the latent space. This result provides an explanation for the superior performance in the programmatic space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchNicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka MarttinenNeurIPS 2024 · 被引用 49 次
- Hierarchical Programmatic Option FrameworkYu-An Lin, Chen-Tao Lee, Chih-Han Yang, Guan-Ting Liu 等NeurIPS 2024 · 被引用 7 次
- Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided SearchMax Liu, Chan-Hung Yu, Wei-Hsu Lee, Cheng-Wei Hung 等ICLR 2025
它引用的顶会 Paper10
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- BUSTLE: Bottom-Up Program Synthesis Through Learning-Guided ExplorationAugustus Odena, Kensen Shi, David Bieber, Rishabh Singh 等ICLR 2021 · 被引用 60 次
- Leveraging Language to Learn Program Abstractions and Search HeuristicsCatherine Wong, Kevin Ellis, Joshua B. Tenenbaum, Jacob AndreasICML 2021 · 被引用 59 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 被引用 42 次
相关 Paper
- Hierarchical Programmatic Reinforcement Learning via Learning to Compose ProgramsGuan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-Yi Lee 等ICML 2023 · 被引用 21 次
- Learning Differentiable Programs with Admissible Neural HeuristicsAmeesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma 等NeurIPS 2020 · 被引用 56 次
- Meta-learning curiosity algorithmsFerran Alet, Martin F. Schneider, Tomás Lozano-Pérez, Leslie Pack KaelblingICLR 2020 · 被引用 67 次
- Revisiting OOD Generalization in Programmatic RLAmirhossein Rajabpour, Kiarash Aghakasiri, Sandra Zilles, Levi LelisICML 2026
- RLang: A Declarative Language for Describing Partial World Knowledge to Reinforcement Learning AgentsRafael Rodríguez-Sánchez, Benjamin Adin Spiegel, Jennifer Wang, Roma Patel 等ICML 2023 · 被引用 4 次
