The Value Equivalence Principle for Model-Based Reinforcement Learning
Christopher Grimm, André Barreto, Satinder Singh, David Silver
摘要
Learning models of the environment from data is often viewed as an essential component to building intelligent reinforcement learning (RL) agents. The common practice is to separate the learning of the model from its use, by constructing a model of the environment's dynamics that correctly predicts the observed state transitions. In this paper we argue that the limited representational resources of model-based RL agents are better used to build models that are directly useful for value-based planning. As our main contribution, we introduce the principle of value equivalence: two models are value equivalent with respect to a set of functions and policies if they yield the same Bellman updates. We propose a formulation of the model learning problem based on the value equivalence principle and analyze how the set of feasible solutions is impacted by the choice of policies and functions. Specifically, we show that, as we augment the set of policies and functions considered, the class of value equivalent models shrinks, until eventually collapsing to a single point corresponding to a model that perfectly describes the environment. In many problems, directly modelling state-to-state transitions may be both difficult and unnecessary. By leveraging the value-equivalence principle one may find simpler models without compromising performance, saving computation and memory. We illustrate the benefits of value-equivalent model learning with experiments comparing it against more traditional counterparts like maximum likelihood estimation. More generally, we argue that the principle of value equivalence underlies a number of recent empirical successes in RL, such as Value Iteration Networks, the Predictron, Value Prediction Networks, TreeQN, and MuZero, and provides a first theoretical underpinning of those results. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert 等ICLR 2022 · 被引用 79 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez 等ICML 2021 · 被引用 69 次
- Towards Robust Bisimulation Metric LearningMete Kemertas, Tristan Aumentado-ArmstrongNeurIPS 2021 · 被引用 68 次
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RLBenjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 57 次
它引用的顶会 Paper3
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 被引用 171 次
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi 等AAAI 2021 · 被引用 76 次
相关 Paper
- Proper Value EquivalenceChristopher Grimm, André Barreto, Gregory Farquhar, David Silver 等NeurIPS 2021 · 被引用 49 次
- Approximate Value EquivalenceChristopher Grimm, André Barreto, Satinder SinghNeurIPS 2022 · 被引用 7 次
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 被引用 25 次
- Calibrated Value-Aware Model Learning with Probabilistic Environment ModelsClaas Voelcker, Anastasiia Pedan, Arash Ahmadian, Romina Abachi 等ICML 2025
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
