Universal Value-Function Uncertainties
Moritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart, Caroline Horsch, Wendelin Boehmer, Matthijs T. J. Spaan
Abstract
Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, and offline RL. While deep ensembles provide a robust method for quantifying value uncertainty, they come with significant computational overhead. Single-model methods, while computationally favorable, often rely on heuristics and typically require additional propagation mechanisms for myopic uncertainty estimates. In this work we introduce universal value-function uncertainties (UVU), which, similar in spirit to random network distillation (RND), quantify uncertainty as squared prediction errors between an online learner and a fixed, randomly initialized target network. Unlike RND, UVU errors reflect policy-conditional , incorporating the future uncertainties may encounter. This is due to the training procedure employed in UVU: the online network is trained using temporal difference learning with a synthetic reward derived from the fixed, randomly initialized target network. We provide an extensive theoretical analysis of our approach using neural tangent kernel (NTK) theory and show that in the limit of infinite network width, UVU errors are exactly equivalent to the variance of an ensemble of independent universal value functions. Empirically, we show that UVU achieves equal performance to large ensembles on challenging multi-task offline RL settings, while offering simplicity and substantial computational savings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfb30e12-4325-48de-8cb4-18bc71754547Builds on27
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 430 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
Related papers
- Contextual Similarity Distillation: Ensemble Uncertainties with a Single ModelMoritz Akiya Zanger, Pascal R. Van der Vaart, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2026 · 6 citations
- Single Model Uncertainty Estimation via Stochastic Data CenteringJayaraman J. Thiagarajan, Rushil Anirudh, Vivek Sivaraman Narayanaswamy, Timo BremerNeurIPS 2022 · 35 citations
- Disentangling the Predictive Variance of Deep Ensembles through the Neural Tangent KernelSeijin Kobayashi, Pau Vilimelis Aceituno, Johannes von OswaldNeurIPS 2022 · 4 citations
- Parameterized Indexed Value Function for Efficient Exploration in Reinforcement LearningTian Tan, Zhihan Xiong, Vikranth R. DwaracherlaAAAI 2020 · 5 citations
- Uncertainty Quantification with the Empirical Neural Tangent KernelJoseph Wilson, Chris van der Heide, Liam Hodgkinson, Fred RoostaNeurIPS 2025 · 11 citations
