The Neural Testbed: Evaluating Joint Predictions
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Dieterich Lawson, Botao Hao, Brendan O'Donoghue, Benjamin Van Roy
Abstract
Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their marginal predictions per input, but also on their joint predictions across many inputs. We evaluate a range of agents using a simple neural network data generating process. Our results indicate that some popular Bayesian deep learning agents do not fare well with joint predictions, even when they can produce accurate marginal predictions. We also show that the quality of joint predictions drives performance in downstream decision tasks. We find these results are robust across choice a wide range of generative models, and highlight the practical importance of joint predictions to the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7b2c3a4-aa51-443a-8553-38c22e1582b4Cited by top-tier papers12
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla et al.NeurIPS 2023 · 142 citations
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum et al.ICML 2022 · 83 citations
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak et al.ICML 2024 · 44 citations
- An Analysis of Ensemble SamplingChao Qin, Zheng Wen, Xiuyuan Lu, Benjamin Van RoyNeurIPS 2022 · 30 citations
- Leveraging Demonstrations to Improve Online Learning: Quality MattersBotao Hao, Rahul Jain, Tor Lattimore, Benjamin Van Roy et al.ICML 2023 · 13 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla et al.NeurIPS 2023 · 142 citations
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 136 citations
- Hypermodels for ExplorationVikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband et al.ICLR 2020 · 49 citations
Related papers
- Tractable Function-Space Variational Inference in Bayesian Neural NetworksTim G. J. Rudner, Zonghao Chen, Yee Whye Teh, Yarin GalNeurIPS 2022 · 70 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Uncertainty Quantification for Deep Regression using Contextualised Normalizing FlowsAdriel Sosa Marco, John Daniel Kirwan, Alexia Toumpa, Simos GerasimouNeurIPS 2025 · 4 citations
- A Hierarchical Variational Neural Uncertainty Model for Stochastic Video PredictionMoitreya Chatterjee, Narendra Ahuja, Anoop CherianICCV 2021 · 18 citations
- Quantifying Uncertainty in the Presence of Distribution ShiftsYuli Slavutsky, David M. BleiNeurIPS 2025 · 2 citations
