The Neural Testbed: Evaluating Joint Predictions
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Dieterich Lawson, Botao Hao, Brendan O'Donoghue, Benjamin Van Roy
摘要
Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their marginal predictions per input, but also on their joint predictions across many inputs. We evaluate a range of agents using a simple neural network data generating process. Our results indicate that some popular Bayesian deep learning agents do not fare well with joint predictions, even when they can produce accurate marginal predictions. We also show that the quality of joint predictions drives performance in downstream decision tasks. We find these results are robust across choice a wide range of generative models, and highlight the practical importance of joint predictions to the community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla 等NeurIPS 2023 · 被引用 142 次
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum 等ICML 2022 · 被引用 83 次
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak 等ICML 2024 · 被引用 44 次
- An Analysis of Ensemble SamplingChao Qin, Zheng Wen, Xiuyuan Lu, Benjamin Van RoyNeurIPS 2022 · 被引用 30 次
- Leveraging Demonstrations to Improve Online Learning: Quality MattersBotao Hao, Rahul Jain, Tor Lattimore, Benjamin Van Roy 等ICML 2023 · 被引用 13 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla 等NeurIPS 2023 · 被引用 142 次
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 被引用 136 次
- Hypermodels for ExplorationVikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband 等ICLR 2020 · 被引用 49 次
相关 Paper
- Tractable Function-Space Variational Inference in Bayesian Neural NetworksTim G. J. Rudner, Zonghao Chen, Yee Whye Teh, Yarin GalNeurIPS 2022 · 被引用 70 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Uncertainty Quantification for Deep Regression using Contextualised Normalizing FlowsAdriel Sosa Marco, John Daniel Kirwan, Alexia Toumpa, Simos GerasimouNeurIPS 2025 · 被引用 4 次
- A Hierarchical Variational Neural Uncertainty Model for Stochastic Video PredictionMoitreya Chatterjee, Narendra Ahuja, Anoop CherianICCV 2021 · 被引用 18 次
- Quantifying Uncertainty in the Presence of Distribution ShiftsYuli Slavutsky, David M. BleiNeurIPS 2025 · 被引用 2 次
