Constructing an Optimal Behavior Basis for the Option Keyboard
Lucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da Silva
Abstract
Multi-task reinforcement learning aims to quickly identify solutions for new tasks with minimal or no additional interaction with the environment. Generalized Policy Improvement (GPI) addresses this by combining a set of base policies to produce a new one that is at least as good -- though not necessarily optimal -- as any individual base policy. Optimality can be ensured, particularly in the linear-reward case, via techniques that compute a Convex Coverage Set (CCS). However, these are computationally expensive and do not scale to complex domains. The Option Keyboard (OK) improves upon GPI by producing policies that are at least as good -- and often better. It achieves this through a learned meta-policy that dynamically combines base policies. However, its performance critically depends on the choice of base policies. This raises a key question: is there an optimal set of base policies -- an optimal behavior basis -- that enables zero-shot identification of optimal solutions for any linear tasks? We solve this open problem by introducing a novel method that efficiently constructs such an optimal behavior basis. We show that it significantly reduces the number of base policies needed to ensure optimality in new tasks. We also prove that it is strictly more expressive than a CCS, enabling particular classes of non-linear tasks to be solved optimally. We empirically evaluate our technique in challenging domains and show that it outperforms state-of-the-art approaches, increasingly so as task complexity increases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 785a1f00-8e41-4b7d-b0e2-26511efa885dBuilds on15
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus et al.ICLR 2024 · 106 citations
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 36 citations
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 26 citations
- Discovering a set of policies for the worst case rewardTom Zahavy, André Barreto, Daniel J. Mankowitz, Shaobo Hou et al.ICLR 2021 · 26 citations
Related papers
- Combining Behaviors with the Successor Features KeyboardWilka Carvalho, Andre Saraiva, Angelos Filos, Andrew K. Lampinen et al.NeurIPS 2023 · 13 citations
- Constructing a Good Behavior Basis for Transfer using Generalized Policy UpdatesSafa Alver, Doina PrecupICLR 2022 · 19 citations
- Multi-Step Generalized Policy Improvement by Leveraging Approximate ModelsLucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da SilvaNeurIPS 2023 · 7 citations
- Optimistic Task Inference for Behavior Foundation ModelsThomas Rupf, Marco Bagatella, Marin Vlastelica, Andreas KrauseICLR 2026 · 6 citations
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu et al.NeurIPS 2021 · 38 citations
