Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
Beyazit Yalcinkaya, Niklas Lauffer, Marcell Vazquez-Chanlatte, Sanjit A. Seshia
摘要
Goal-conditioned reinforcement learning is a powerful way to control an AI agent's behavior at runtime. That said, popular goal representations, e.g., target states or natural language, are either limited to Markovian tasks or rely on ambiguous task semantics. We propose representing temporal goals using compositions of deterministic finite automata (cDFAs) and use cDFAs to guide RL agents. cDFAs balance the need for formal temporal semantics with ease of interpretation: if one can understand a flow chart, one can understand a cDFA. On the other hand, cDFAs form a countably infinite concept class with Boolean semantics, and subtle changes to the automaton can result in very different tasks, making them difficult to condition agent behavior on. To address this, we observe that all paths through a DFA correspond to a series of reach-avoid tasks and propose pre-training graph neural network embeddings on"reach-avoid derived"DFAs. Through empirical evaluation, we demonstrate that the proposed pre-training method enables zero-shot generalization to various cDFA task classes and accelerated policy specialization without the myopic suboptimality of hierarchical methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement LearningZijian Guo, Ilker Isik, H. M. Sabbir Ahmad, Wenchao LiNeurIPS 2025 · 被引用 13 次
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement LearningRoger Creus Castanyer, Faisal Mohamed, Pablo Samuel Castro, Cyrus Neary 等ICLR 2026 · 被引用 6 次
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari 等NeurIPS 2025 · 被引用 5 次
- Automaton Constrained Q-LearningAnastasios Manganaris, Vittorio Giammarino, Ahmed H. QureshiNeurIPS 2025 · 被引用 3 次
- Accelerated Learning with Linear Temporal Logic using Differentiable SimulationAlper Kamil Bozkurt, Calin Belta, Ming C. LinICLR 2026 · 被引用 2 次
它引用的顶会 Paper7
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch 等NeurIPS 2023 · 被引用 182 次
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 被引用 112 次
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 被引用 106 次
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
相关 Paper
- Compositional Policy Learning in Stochastic Control Systems with Formal GuaranteesDorde Zikelic, Mathias Lechner, Abhinav Verma, Krishnendu Chatterjee 等NeurIPS 2023 · 被引用 31 次
- Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement LearningRiashat Islam, Hongyu Zang, Anirudh Goyal, Alex Lamb 等NeurIPS 2022 · 被引用 10 次
- In a Nutshell, the Human Asked for This: Latent Goals for Following Temporal SpecificationsBorja G. León, Murray Shanahan, Francesco BelardinelliICLR 2022 · 被引用 23 次
- Pretraining Representations for Data-Efficient Reinforcement LearningMax Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand 等NeurIPS 2021 · 被引用 151 次
- Imitation Learning with Temporal Logic ConstraintsZining Fan, He ZhuNeurIPS 2025 · 被引用 2 次
