Approximate Value Equivalence
Christopher Grimm, André Barreto, Satinder Singh
摘要
Model-based reinforcement learning agents must make compromises about which aspects of the environment their models should capture. The value equivalence (VE) principle posits that these compromises should be made considering the model’s eventual use in value-based planning. Given sets of functions and policies, a model is said to be order-k VE to the environment if k applications of the Bellman operators induced by the policies produce the correct result when applied to the functions. Prior work investigated the classes of models induced by VE when we vary k and the sets of policies and functions. This gives rise to a rich collection of topological relationships and conditions under which VE models are optimal for planning. Despite this effort, relatively little is known about the planning performance of models that fail to satisfy these conditions. This is due to the rigidity of the VE formalism, as classes of VE models are defined with respect to exact constraints on their Bellman operators. This limitation gets amplified by the fact that such constraints themselves may depend on functions that can only be approximated in practice. To address these problems we propose approximate value equivalence (AVE), which extends the VE formalism by replacing equalities with error tolerances. This extension allows us to show that AVE models with respect to one set of functions are also AVE with respect to any other set of functions if we tolerate a high enough error. We can then derive bounds on the performance of VE models with respect to arbitrary sets of functions . Moreover, AVE models more accurately reflect what can be learned by our agents in practice, allowing us to investigate previously unexplored tensions between model capacity and the choice of VE model class. In contrast to previous works, we show empirically that there are situations where agents with limited capacity should prefer to learn more accurate models with respect to smaller sets of functions over less accurate models with respect to larger sets of functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Distributional Model Equivalence for Risk-Sensitive Reinforcement LearningTyler Kastner, Murat A. Erdogdu, Amir-massoud FarahmandNeurIPS 2023 · 被引用 9 次
- The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement LearningSarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan A. Murphy 等ICML 2023 · 被引用 6 次
它引用的顶会 Paper5
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
- Proper Value EquivalenceChristopher Grimm, André Barreto, Gregory Farquhar, David Silver 等NeurIPS 2021 · 被引用 49 次
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 被引用 47 次
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 被引用 37 次
- Self-Consistent Models and ValuesGregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos 等NeurIPS 2021 · 被引用 10 次
相关 Paper
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 被引用 25 次
- The Benefits of Model-Based Generalization in Reinforcement LearningKenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen SchmidhuberICML 2023 · 被引用 18 次
- Learning to Execute: Efficient Learning of Universal Plan-Conditioned Policies in RoboticsIngmar Schubert, Danny Driess, Ozgur S. Oguz, Marc ToussaintNeurIPS 2021 · 被引用 2 次
- Maximum Entropy Model Correction in Reinforcement LearningAmin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, Amir-massoud FarahmandICLR 2024 · 被引用 3 次
- Model-Value Inconsistency as a Signal for Epistemic UncertaintyAngelos Filos, Eszter Vértes, Zita Marinho, Gregory Farquhar 等ICML 2022 · 被引用 9 次
