Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
Daniel D. Johnson, Daniel Tarlow, David Duvenaud, Chris J. Maddison
摘要
Identifying how much a model knows about the stochastic real-world process it was trained on is important to ensure it avoids producing incorrect or"hallucinated"answers or taking unsafe actions. But this is difficult for generative models because probabilistic predictions do not distinguish between per-response noise (aleatoric uncertainty) and lack of knowledge about the process (epistemic uncertainty), and existing epistemic uncertainty quantification techniques tend to be overconfident when the model underfits. We propose a general strategy for teaching a model to both approximate and also estimate the remaining gaps between and : train it to predict pairs of independent responses drawn from the true conditional distribution, allow it to"cheat"by observing one response while predicting the other, then measure how much it cheats. Remarkably, we prove that being good at cheating (i.e. cheating whenever it improves your prediction) is equivalent to being second-order calibrated, a principled extension of ordinary calibration that allows us to construct provably-correct frequentist confidence intervals for and detect incorrect responses with high probability. We demonstrate empirically that our approach accurately estimates how much models don't know across ambiguous image classification, (synthetic) language modeling, and partially-observable navigation tasks, outperforming existing techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Estimating the Hallucination Rate of Generative AIAndrew Jesson, Nicolas Beltran-Velez, Quentin Chu, Sweta Karlekar 等NeurIPS 2024 · 被引用 46 次
- Know What You Don't Know: Uncertainty Calibration of Process Reward ModelsYoung-Jin Park, Kristjan Greenewald, Kaveh Alimohammadi, Hao Wang 等NeurIPS 2025 · 被引用 21 次
- Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty QuantificationKimia Hamidieh, Veronika Thost, Walter Gerych, Mikhail Yurochkin 等ICLR 2026 · 被引用 12 次
- Looking Inward: Language Models Can Learn About Themselves by IntrospectionFelix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight 等ICLR 2025 · 被引用 5 次
- Efficient semantic uncertainty quantification in language models via diversity-steered samplingJi Won Park, Kyunghyun ChoNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran 等NeurIPS 2020 · 被引用 604 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
相关 Paper
- Pitfalls of Epistemic Uncertainty Quantification through Loss MinimisationViktor Bengs, Eyke Hüllermeier, Willem WaegemanNeurIPS 2022 · 被引用 78 次
- On Second-Order Scoring Rules for Epistemic Uncertainty QuantificationViktor Bengs, Eyke Hüllermeier, Willem WaegemanICML 2023 · 被引用 37 次
- Epistemic Uncertainty Quantification To Improve Decisions From Black-Box ModelsSébastien Melo, Gaël Varoquaux, Marine Le MorvanICLR 2026
- Provable Uncertainty Decomposition via Higher-Order CalibrationGustaf Ahdritz, Aravind Gollakota, Parikshit Gopalan, Charlotte Peale 等ICLR 2025
- Epistemic Uncertainty Quantification for Pretrained Neural NetworksHanjing Wang, Qiang JiCVPR 2024 · 被引用 5 次
