U-trustworthy Models. Reliability, Competence, and Confidence in Decision-Making
Ritwik Vashistha, Arya Farahi
摘要
With growing concerns regarding bias and discrimination in predictive models, the AI community has increasingly focused on assessing AI system trustworthiness. Conventionally, trustworthy AI literature relies on the probabilistic framework and calibration as prerequisites for trustworthiness. In this work, we depart from this viewpoint by proposing a novel trust framework inspired by the philosophy literature on trust. We present a precise mathematical definition of trustworthiness, termed U -trustworthiness, specifically tailored for a subset of tasks aimed at maximizing a utility function. We argue that a model's U-trustworthiness is contingent upon its ability to maximize Bayes utility within this task subset. Our first set of results challenges the probabilistic framework by demonstrating its potential to favor less trustworthy models and introduce the risk of misleading trustworthiness assessments. Within the context of Utrustworthiness, we prove that properly-ranked models are inherently U -trustworthy. Furthermore, we advocate for the adoption of the AUC metric as the preferred measure of trustworthiness. By offering both theoretical guarantees and experimental validation, AUC enables robust evaluation of trustworthiness, thereby enhancing model selection and hyperparameter tuning to yield more trustworthy outcomes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Towards Trustworthy Predictions from Deep Neural Networks with Fast Adversarial CalibrationChristian Tomani, Florian BuettnerAAAI 2021 · 被引用 43 次
- The Value of AI Guidance in Human Examination of Synthetically-Generated FacesAidan Boyd, Patrick Tinsley, Kevin W. Bowyer, Adam CzajkaAAAI 2023 · 被引用 20 次
- Learning to Predict Trustworthiness with Steep Slope LossYan Luo, Yongkang Wong, Mohan S. Kankanhalli, Qi ZhaoNeurIPS 2021 · 被引用 17 次
- Evaluating the Calibration of Knowledge Graph Embeddings for Trustworthy Link PredictionTara Safavi, Danai Koutra, Edgar MeijEMNLP 2020 · 被引用 1 次
相关 Paper
- Minimax AUC Fairness: Efficient Algorithm with Provable ConvergenceZhenhuan Yang, Yan Lok Ko, Kush R. Varshney, Yiming YingAAAI 2023 · 被引用 22 次
- A Closer Look at AUROC and AUPRC under Class ImbalanceMatthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti 等NeurIPS 2024 · 被引用 191 次
- The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and TrustNishant Subramani, Palash Goyal, Yiwen Song, Mani Malek 等ICML 2026 · 被引用 1 次
- Overcoming Common Flaws in the Evaluation of Selective Classification SystemsJeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner 等NeurIPS 2024 · 被引用 44 次
- AUC Optimization with a Reject OptionSong-Qing Shen, Bin-Bin Yang, Wei GaoAAAI 2020 · 被引用 7 次
