Exploration via Epistemic Value Estimation
Simon Schmitt, John Shawe-Taylor, Hado van Hasselt
摘要
How to efficiently explore in reinforcement learning is an open problem. Many exploration algorithms employ the epistemic uncertainty of their own value predictions -- for instance to compute an exploration bonus or upper confidence bound. Unfortunately the required uncertainty is difficult to estimate in general with function approximation.
We propose epistemic value estimation (EVE): a recipe that is compatible with sequential decision making and with neural network function approximators. It equips agents with a tractable posterior over all their parameters from which epistemic value uncertainty can be computed efficiently.
We use the recipe to derive an epistemic Q-Learning agent and observe competitive performance on a series of benchmarks. Experiments confirm that the EVE recipe facilitates efficient exploration in hard exploration tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bayesian Exploration NetworksMattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon WhitesonICML 2024 · 被引用 4 次
- Epistemic Bellman OperatorsPascal R. van der Vaart, Matthijs T. J. Spaan, Neil Yorke-SmithAAAI 2025 · 被引用 2 次
- Universal Value-Function UncertaintiesMoritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper2
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
相关 Paper
- Efficient Exploration for LLMsVikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, Benjamin Van RoyICML 2024 · 被引用 45 次
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao 等ICML 2021 · 被引用 46 次
- Implicit Generative Modeling for Efficient ExplorationNeale Ratzlaff, Qinxun Bai, Fuxin Li, Wei XuICML 2020 · 被引用 15 次
- EUBRL: Epistemic Uncertainty Directed Bayesian Reinforcement LearningJianfei Ma, Wee Sun LeeICLR 2026 · 被引用 2 次
- Model-Value Inconsistency as a Signal for Epistemic UncertaintyAngelos Filos, Eszter Vértes, Zita Marinho, Gregory Farquhar 等ICML 2022 · 被引用 9 次
