Multi-Step Generalized Policy Improvement by Leveraging Approximate Models
Lucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da Silva
Abstract
We introduce a principled method for performing zero-shot transfer in reinforcement learning (RL) by exploiting approximate models of the environment. Zeroshot transfer in RL has been investigated by leveraging methods rooted in generalized policy improvement (GPI) and successor features (SFs). Although computationally efficient, these methods are model-free: they analyze a library of policies-each solving a particular task-and identify which action the agent should take. We investigate the more general setting where, in addition to a library of policies, the agent has access to an approximate environment model. Even though model-based RL algorithms can identify near-optimal policies, they are typically computationally intensive. We introduce h-GPI, a multi-step extension of GPI that interpolates between these extremes-standard model-free GPI and fully model-based planning-as a function of a parameter, h, regulating the amount of time the agent has to reason. We prove that h-GPI's performance lower bound is strictly better than GPI's, and show that h-GPI generally outperforms GPI as h increases. Furthermore, we prove that as h increases, h-GPI's performance becomes arbitrarily less susceptible to sub-optimality in the agent's policy library. Finally, we introduce novel bounds characterizing the gains achievable by h-GPI as a function of approximation errors in both the agent's policy library and its (possibly learned) model. These bounds strictly generalize those known in the literature. We evaluate h-GPI on challenging tabular and continuous-state problems under value function approximation and show that it consistently outperforms GPI and state-of-the-art competing methods under various levels of approximation errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 7 citations
- Constructing an Optimal Behavior Basis for the Option KeyboardLucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da SilvaNeurIPS 2025 · 4 citations
Builds on20
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley et al.ICLR 2020 · 176 citations
Related papers
- Constrained GPI for Zero-Shot Transfer in Reinforcement LearningJaekyeom Kim, Seohong Park, Gunhee KimNeurIPS 2022 · 10 citations
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du et al.NeurIPS 2024 · 11 citations
- SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement LearningShuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan et al.ICML 2024 · 7 citations
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney et al.ICML 2022 · 11 citations
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee et al.ICLR 2023 · 2 citations
