Great Models Think Alike: Improving Model Reliability via Inter-Model Latent Agreement
Ailin Deng, Miao Xiong, Bryan Hooi
摘要
Reliable application of machine learning is of primary importance to the practical deployment of deep learning methods. A fundamental challenge is that models are often unreliable due to overconfidence (Hendrycks & Gimpel, 2017) . In this paper, we estimate a model's reliability by measuring the agreement between its latent space, and the latent space of a foundation model. However, it is challenging to measure the agreement between two different latent spaces due to their incoherence, e.g., arbitrary rotations and different dimensionality. To overcome this incoherence issue, we design a neighborhood agreement measure between latent spaces and find that this agreement is surprisingly well-correlated with the reliability of a model's predictions. Further, we show that fusing neighborhood agreement into a model's predictive confidence in a post-hoc way significantly improves its reliability. Theoretical analysis and extensive experiments on failure detection across various datasets verify the effectiveness of our method on both in-distribution and out-of-distribution settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- SteerConf: Steering LLMs for Confidence ElicitationZiang Zhou, Tianyuan Jin, Jieming Shi, Qing LiNeurIPS 2025 · 被引用 23 次
- Taming Overconfidence in LLMs: Reward Calibration in RLHFJixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin HuangICLR 2025
- CaTS: Calibrated Test-Time Scaling for Efficient LLM ReasoningChengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu 等ICLR 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Evaluating Deep Neural Networks in Deployment: A Comparative Study (Replicability Study)Eduard Pinconschi, Divya Gopinath, Rui Abreu, Corina S. PasareanuISSTA 2024
- Consistent Counterfactuals for Deep ModelsEmily Black, Zifan Wang, Matt FredriksonICLR 2022 · 被引用 56 次
- Identifying Metric Structures of Deep Latent Variable ModelsStas Syrota, Yevgen Zainchkovskyy, Johnny Xi, Benjamin Bloem-Reddy 等ICML 2025
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
- Combining Statistical Depth and Fermat Distance for Uncertainty QuantificationHai-Vy Nguyen, Fabrice Gamboa, Reda Chhaibi, Sixin Zhang 等NeurIPS 2024 · 被引用 3 次
