Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
Daniel D. Johnson, Daniel Tarlow, David Duvenaud, Chris J. Maddison
Abstract
Identifying how much a model knows about the stochastic real-world process it was trained on is important to ensure it avoids producing incorrect or"hallucinated"answers or taking unsafe actions. But this is difficult for generative models because probabilistic predictions do not distinguish between per-response noise (aleatoric uncertainty) and lack of knowledge about the process (epistemic uncertainty), and existing epistemic uncertainty quantification techniques tend to be overconfident when the model underfits. We propose a general strategy for teaching a model to both approximate and also estimate the remaining gaps between and : train it to predict pairs of independent responses drawn from the true conditional distribution, allow it to"cheat"by observing one response while predicting the other, then measure how much it cheats. Remarkably, we prove that being good at cheating (i.e. cheating whenever it improves your prediction) is equivalent to being second-order calibrated, a principled extension of ordinary calibration that allows us to construct provably-correct frequentist confidence intervals for and detect incorrect responses with high probability. We demonstrate empirically that our approach accurately estimates how much models don't know across ambiguous image classification, (synthetic) language modeling, and partially-observable navigation tasks, outperforming existing techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e113e5ba-c0e1-4274-a41a-809a0e249e13Cited by top-tier papers9
- Estimating the Hallucination Rate of Generative AIAndrew Jesson, Nicolas Beltran-Velez, Quentin Chu, Sweta Karlekar et al.NeurIPS 2024 · 46 citations
- Know What You Don't Know: Uncertainty Calibration of Process Reward ModelsYoung-Jin Park, Kristjan Greenewald, Kaveh Alimohammadi, Hao Wang et al.NeurIPS 2025 · 21 citations
- Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty QuantificationKimia Hamidieh, Veronika Thost, Walter Gerych, Mikhail Yurochkin et al.ICLR 2026 · 12 citations
- Looking Inward: Language Models Can Learn About Themselves by IntrospectionFelix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight et al.ICLR 2025 · 5 citations
- Efficient semantic uncertainty quantification in language models via diversity-steered samplingJi Won Park, Kyunghyun ChoNeurIPS 2025 · 3 citations
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
Related papers
- Pitfalls of Epistemic Uncertainty Quantification through Loss MinimisationViktor Bengs, Eyke Hüllermeier, Willem WaegemanNeurIPS 2022 · 78 citations
- On Second-Order Scoring Rules for Epistemic Uncertainty QuantificationViktor Bengs, Eyke Hüllermeier, Willem WaegemanICML 2023 · 37 citations
- Epistemic Uncertainty Quantification To Improve Decisions From Black-Box ModelsSébastien Melo, Gaël Varoquaux, Marine Le MorvanICLR 2026
- Provable Uncertainty Decomposition via Higher-Order CalibrationGustaf Ahdritz, Aravind Gollakota, Parikshit Gopalan, Charlotte Peale et al.ICLR 2025
- Epistemic Uncertainty Quantification for Pretrained Neural NetworksHanjing Wang, Qiang JiCVPR 2024 · 5 citations
