Model-Value Inconsistency as a Signal for Epistemic Uncertainty
Angelos Filos, Eszter Vértes, Zita Marinho, Gregory Farquhar, Diana Borsa, Abram L. Friesen, Feryal M. P. Behbahani, Tom Schaul, André Barreto, Simon Osindero
Abstract
Using a model of the environment and a value function, an agent can construct many estimates of a state's value, by unrolling the model for different lengths and bootstrapping with its value function. Our key insight is that one can treat this set of value estimates as a type of ensemble, which we call an implicit value ensemble (IVE). Consequently, the discrepancy between these estimates can be used as a proxy for the agent's epistemic uncertainty; we term this signal model-value inconsistency or self-inconsistency for short. Unlike prior work which estimates uncertainty by training an ensemble of many models and/or value functions, this approach requires only the single model and value function which are already being learned in most model-based reinforcement learning algorithms. We provide empirical evidence in both tabular and function approximation settings from pixels that self-inconsistency is useful (i) as a signal for exploration, (ii) for acting safely under distribution shifts, and (iii) for robustifying value-based planning with a learned model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae399909-0ca2-4762-b1cc-ff9a04ed6e8cCited by top-tier papers7
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang et al.NeurIPS 2023 · 41 citations
- Learning General World Models in a Handful of Reward-Free DeploymentsYingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball et al.NeurIPS 2022 · 16 citations
- ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive AdvantagesAndrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez et al.ICML 2024 · 11 citations
- Multi-Step Generalized Policy Improvement by Leveraging Approximate ModelsLucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da SilvaNeurIPS 2023 · 7 citations
Builds on19
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
Related papers
- Self-Consistent Models and ValuesGregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos et al.NeurIPS 2021 · 10 citations
- Perceiving the Knowledge Boundary: Uncertainty-Guided Exploration and Imagination for World ModelsZhenxian Liu, Peixi Peng, Yangru Huang, Yonghong TianAAAI 2026
- Universal Value-Function UncertaintiesMoritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart et al.ICLR 2026 · 1 citation
- Selective Dyna-Style Planning Under Limited Model CapacityZaheer Abbas, Samuel Sokota, Erin Talvitie, Martha WhiteICML 2020 · 38 citations
- One Risk to Rule Them All: A Risk-Sensitive Perspective on Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2023 · 26 citations
