A Boolean Task Algebra for Reinforcement Learning
Geraud Nangue Tasse, Steven James, Benjamin Rosman
摘要
The ability to compose learned skills to solve new tasks is an important property of lifelong-learning agents. In this work, we formalise the logical composition of tasks as a Boolean algebra. This allows us to formulate new tasks in terms of the negation, disjunction and conjunction of a set of base tasks. We then show that by learning goal-oriented value functions and restricting the transition dynamics of the tasks, an agent can solve these new tasks with no further learning. We prove that by composing these value functions in specific ways, we immediately recover the optimal policies for all tasks expressible under the Boolean algebra. We verify our approach in two domains-including a high-dimensional video game environment requiring function approximation-where an agent first learns a set of base skills, and then composes them to solve a super-exponential number of new tasks. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- MoCoDA: Model-based Counterfactual Data AugmentationSilviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh GargNeurIPS 2022 · 被引用 60 次
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
相关 Paper
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 被引用 23 次
- Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningGeraud Nangue Tasse, Devon Jarvis, Steven James, Benjamin RosmanICLR 2024 · 被引用 12 次
- Lifelong Learning of Compositional StructuresJorge A. Mendez, Eric EatonICLR 2021 · 被引用 50 次
- Run-Time Task Composition with Safety SemanticsKevin Leahy, Makai Mann, Zachary SerlinICML 2024 · 被引用 1 次
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without ForgettingJorge A. Mendez, Boyu Wang, Eric EatonNeurIPS 2020 · 被引用 42 次
