Generalised Policy Improvement with Geometric Policy Composition
Shantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney, Rémi Munos, André Barreto
Abstract
We introduce a method for policy improvement that interpolates between the greedy approach of value-based reinforcement learning (RL) and the full planning approach typical of model-based RL. The new method builds on the concept of a geometric horizon model (GHM, also known as a gamma-model), which models the discounted state-visitation distribution of a given policy. We show that we can evaluate any non-Markov policy that switches between a set of base Markov policies with fixed probability by a careful composition of the base policy GHMs, without any additional learning. We can then apply generalised policy improvement (GPI) to collections of such non-Markov policies to obtain a new Markov policy that will in general outperform its precursors. We provide a thorough theoretical analysis of this approach, develop applications to transfer and standard RL, and empirically demonstrate its effectiveness over standard GPI on a challenging deep RL continuous control task. We also provide an analysis of GHM training methods, proving a novel convergence result regarding previously proposed methods and showing how to train these models stably in deep RL settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 159cf3a0-9547-42ae-b142-0f3e17231cebCited by top-tier papers11
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 76 citations
- Bootstrapped Representations in Reinforcement LearningCharline Le Lan, Stephen Tu, Mark Rowland, Anna Harutyunyan et al.ICML 2023 · 12 citations
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang et al.ICML 2024 · 11 citations
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 9 citations
- Multi-Step Generalized Policy Improvement by Leveraging Approximate ModelsLucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da SilvaNeurIPS 2023 · 7 citations
Builds on11
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley et al.ICLR 2020 · 176 citations
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
- Gamma-Models: Generative Temporal Difference Learning for Infinite-Horizon PredictionMichael Janner, Igor Mordatch, Sergey LevineNeurIPS 2020 · 49 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
- Discovering a set of policies for the worst case rewardTom Zahavy, André Barreto, Daniel J. Mankowitz, Shaobo Hou et al.ICLR 2021 · 26 citations
Related papers
- Offline Reinforcement Learning with Universal Horizon ModelsHojun Chung, Junseo Lee, Songhwai OhICML 2026 · 1 citation
- Temporal Difference FlowsJesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Rémi Munos et al.ICML 2025
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 4 citations
- SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement LearningShuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan et al.ICML 2024 · 7 citations
- Geometric Policy Iteration for Markov Decision ProcessesYue Wu, Jesús A. De LoeraKDD 2022 · 1 citation
