Generalisation in Lifelong Reinforcement Learning through Logical Composition
Geraud Nangue Tasse, Steven James, Benjamin Rosman
Abstract
We leverage logical composition in reinforcement learning to create a framework that enables an agent to autonomously determine whether a new task can be immediately solved using its existing abilities, or whether a task-specific skill should be learned. In the latter case, the proposed algorithm also enables the agent to learn the new task faster by generating an estimate of the optimal policy. Importantly, we provide two main theoretical results: we bound the performance of the transferred policy on a new task, and we give bounds on the necessary and sufficient number of tasks that need to be learned throughout an agent's lifetime to generalise over a distribution. We verify our approach in a series of experiments, where we perform transfer learning both after learning a set of base tasks, and after learning an arbitrary set of tasks. We also demonstrate that, as a side effect of our transfer learning approach, an agent can produce an interpretable Boolean expression of its understanding of the current task. Finally, we demonstrate our approach in the full lifelong setting where an agent receives tasks from an unknown distribution. Starting from scratch, an agent is able to quickly generalise over the task distribution after learning only a few tasks, which are sub-logarithmic in the size of the task space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68ef698d-f6bd-463d-b9d1-10071de9dfc7Cited by top-tier papers6
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 36 citations
- Constructing a Good Behavior Basis for Transfer using Generalized Policy UpdatesSafa Alver, Doina PrecupICLR 2022 · 19 citations
- Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement LearningJacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. KulkarniAAAI 2023 · 13 citations
- Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningGeraud Nangue Tasse, Devon Jarvis, Steven James, Benjamin RosmanICLR 2024 · 12 citations
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli et al.ICLR 2023 · 2 citations
Builds on4
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu et al.ICLR 2021 · 161 citations
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 126 citations
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 71 citations
- Constructing a Good Behavior Basis for Transfer using Generalized Policy UpdatesSafa Alver, Doina PrecupICLR 2022 · 19 citations
Related papers
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson et al.ICLR 2020 · 35 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 106 citations
- The Logical Options FrameworkBrandon Araki, Xiao Li, Kiran Vodrahalli, Jonathan A. DeCastro et al.ICML 2021 · 44 citations
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui et al.ICLR 2026 · 46 citations
