Neurosymbolic World Models for Sequential Decision Making
Leonardo Hernandez Cano, Maxine Perroni-Scharf, Neil Dhir, Arun Ramamurthy, Armando Solar-Lezama
摘要
We present Structured World Modeling for Policy Optimization (SWMPO), a framework for unsupervised learning of neurosymbolic Finite State Machines (FSM) that capture environmental structure for policy optimization. SWMPO models the environment as a FSM, where each state corresponds to a specific region of the state space with distinct dynamics (e.g., water and land). This structured representation can be leveraged for tasks like policy optimization. Our proposed FSM synthesis algorithm operates in an unsupervised manner, leveraging low-level features from unprocessed, non-visual data to learn non-linear models, making it adaptable across various domains. The synthesized FSM models are expressive enough to be used in a model-based Reinforcement Learning scheme that leverages offline data to efficiently synthesize environment-specific world models. We demonstrate the advantages of SWMPO by benchmarking its environment modeling capabilities in a number of simulation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu 等NeurIPS 2020 · 被引用 736 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Recurrent Switching Dynamical Systems Models for Multiple Interacting Neural PopulationsJoshua I. Glaser, Matthew R. Whiteway, John P. Cunningham, Liam Paninski 等NeurIPS 2020 · 被引用 113 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- Hidden Parameter Recurrent State Space Models For Changing Dynamics ScenariosVaisakh Shaj, Dieter Büchler, Rohit Sonker, Philipp Becker 等ICLR 2022 · 被引用 9 次
相关 Paper
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller 等NeurIPS 2025 · 被引用 31 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual ExplanationsLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang 等ICML 2024 · 被引用 18 次
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied ReasoningWonje Choi, Jooyoung Kim, Honguk WooNeurIPS 2025 · 被引用 4 次
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari 等NeurIPS 2025 · 被引用 5 次
