Induction of Subgoal Automata for Reinforcement Learning
Daniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda, Anders Jonsson
摘要
In this work we present ISA, a novel approach for learning and exploiting subgoals in reinforcement learning (RL). Our method relies on inducing an automaton whose transitions are subgoals expressed as propositional formulas over a set of observable events. A state-of-the-art inductive logic programming system is used to learn the automaton from observation traces perceived by the RL agent. The reinforcement learning and automaton learning processes are interleaved: a new refined automaton is learned whenever the RL agent generates a trace not recognized by the current automaton. We evaluate ISA in several gridworld problems and show that it performs similarly to a method for which automata are given in advance. We also show that the learned automata can be exploited to speed up convergence through reward shaping and transfer learning across multiple tasks. Finally, we analyze the running time and the number of traces that ISA needs to learn an automata, and the impact that the number of observable events has on the learner's performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham 等AAAI 2021 · 被引用 62 次
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu 等AAAI 2021 · 被引用 40 次
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 被引用 37 次
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor 等NeurIPS 2024 · 被引用 19 次
相关 Paper
- EAGER: Asking and Answering Questions for Automatic Reward Shaping in Language-guided RLThomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain LamprierNeurIPS 2022 · 被引用 35 次
- Imitating Graph-Based Planning with Goal-Conditioned PoliciesJunsu Kim, Younggyo Seo, Sungsoo Ahn, Kyunghwan Son 等ICLR 2023 · 被引用 2 次
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah 等AAAI 2021 · 被引用 54 次
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 被引用 23 次
