Beyond Confidence: Reliable Models Should Also Consider Atypicality
Mert Yüksekgönül, Linjun Zhang, James Y. Zou, Carlos Guestrin
摘要
While most machine learning models can provide confidence in their predictions, confidence is insufficient to understand a prediction's reliability. For instance, the model may have a low confidence prediction if the input is not well-represented in the training dataset or if the input is inherently ambiguous. In this work, we investigate the relationship between how atypical (rare) a sample or a class is and the reliability of a model's predictions. We first demonstrate that atypicality is strongly related to miscalibration and accuracy. In particular, we empirically show that predictions for atypical inputs or atypical classes are more overconfident and have lower accuracy. Using these insights, we show incorporating atypicality improves uncertainty quantification and model performance for discriminative neural networks and large language models. In a case study, we show that using atypicality improves the performance of a skin lesion classifier across different skin tone groups without having access to the group attributes. Overall, we propose that models should use not only confidence but also atypicality to improve uncertainty quantification and performance. Our results demonstrate that simple post-hoc atypicality estimators can provide significant value. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language ModelsMert Yüksekgönül, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar 等ICLR 2024 · 被引用 73 次
- When is Multicalibration Post-Processing Necessary?Dutch Hansen, Siddartha Devic, Preetum Nakkiran, Vatsal SharanNeurIPS 2024 · 被引用 21 次
- Open-Vocabulary Calibration for Fine-tuned CLIPShuoyuan Wang, Jindong Wang, Guoqing Wang, Bob Zhang 等ICML 2024 · 被引用 17 次
- Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AINabeel Seedat, Fergus Imrie, Mihaela van der SchaarICLR 2024 · 被引用 16 次
- Beyond probability partitions: Calibrating neural networks with semantic aware groupingJia-Qi Yang, De-Chuan Zhan, Le GanNeurIPS 2023 · 被引用 14 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran 等NeurIPS 2020 · 被引用 604 次
相关 Paper
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 被引用 177 次
- Calibrating Predictions to Decisions: A Novel Approach to Multi-Class CalibrationShengjia Zhao, Michael P. Kim, Roshni Sahoo, Tengyu Ma 等NeurIPS 2021 · 被引用 96 次
- Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?Wataru Hashimoto, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 被引用 1 次
- Nonparametric Uncertainty Quantification for Single Deterministic Neural NetworkNikita Kotelevskii, Aleksandr Artemenkov, Kirill Fedyanin, Fedor Noskov 等NeurIPS 2022 · 被引用 52 次
- Learning Sample Difficulty from Pre-trained Models for Reliable PredictionPeng Cui, Dan Zhang, Zhijie Deng, Yinpeng Dong 等NeurIPS 2023 · 被引用 21 次
