Proper Value Equivalence
Christopher Grimm, André Barreto, Gregory Farquhar, David Silver, Satinder Singh
Abstract
One of the main challenges in model-based reinforcement learning (RL) is to decide which aspects of the environment should be modeled. The value-equivalence (VE) principle proposes a simple answer to this question: a model should capture the aspects of the environment that are relevant for value-based planning. Technically, VE distinguishes models based on a set of policies and a set of functions: a model is said to be VE to the environment if the Bellman operators it induces for the policies yield the correct result when applied to the functions. As the number of policies and functions increase, the set of VE models shrinks, eventually collapsing to a single point corresponding to a perfect model. A fundamental question underlying the VE principle is thus how to select the smallest sets of policies and functions that are sufficient for planning. In this paper we take an important step towards answering this question. We start by generalizing the concept of VE to order- counterparts defined with respect to applications of the Bellman operator. This leads to a family of VE classes that increase in size as . In the limit, all functions become value functions, and we have a special instantiation of VE which we call proper VE or simply PVE. Unlike VE, the PVE class may contain multiple models even in the limit when all value functions are used. Crucially, all these models are sufficient for planning, meaning that they will yield an optimal policy despite the fact that they may ignore many aspects of the environment. We construct a loss function for learning PVE models and argue that popular algorithms such as MuZero can be understood as minimizing an upper bound for this loss. We leverage this connection to propose a modification to MuZero and show that it can lead to improved performance in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd01764f-fd45-4ba1-90e0-1722da128f57Cited by top-tier papers14
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 37 citations
- Continuous MDP Homomorphisms and Homomorphic Policy GradientSahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger et al.NeurIPS 2022 · 34 citations
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 25 citations
- Distributional Model Equivalence for Risk-Sensitive Reinforcement LearningTyler Kastner, Murat A. Erdogdu, Amir-massoud FarahmandNeurIPS 2023 · 9 citations
- Model-Value Inconsistency as a Signal for Epistemic UncertaintyAngelos Filos, Eszter Vértes, Zita Marinho, Gregory Farquhar et al.ICML 2022 · 9 citations
Builds on6
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
Related papers
- Approximate Value EquivalenceChristopher Grimm, André Barreto, Satinder SinghNeurIPS 2022 · 7 citations
- Calibrated Value-Aware Model Learning with Probabilistic Environment ModelsClaas Voelcker, Anastasiia Pedan, Arash Ahmadian, Romina Abachi et al.ICML 2025
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert et al.ICLR 2022 · 79 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
