Learning to Predict Trustworthiness with Steep Slope Loss
Yan Luo, Yongkang Wong, Mohan S. Kankanhalli, Qi Zhao
Abstract
Understanding the trustworthiness of a prediction yielded by a classifier is critical for the safe and effective use of AI models. Prior efforts have been proven to be reliable on small-scale datasets. In this work, we study the problem of predicting trustworthiness on real-world large-scale datasets, where the task is more challenging due to high-dimensional features, diverse visual concepts, and a large number of samples. In such a setting, we observe that the trustworthiness predictors trained with prior-art loss functions, i.e., the cross entropy loss, focal loss, and true class probability confidence loss, are prone to view both correct predictions and incorrect predictions to be trustworthy. The reasons are two-fold. Firstly, correct predictions are generally dominant over incorrect predictions. Secondly, due to the data complexity, it is challenging to differentiate the incorrect predictions from the correct ones on real-world large-scale datasets. To improve the generalizability of trustworthiness predictors, we propose a novel steep slope loss to separate the features w.r.t. correct predictions from the ones w.r.t. incorrect predictions by two slide-like curves that oppose each other. The proposed loss is evaluated with two representative deep learning models, i.e., Vision Transformer and ResNet, as trustworthiness predictors. We conduct comprehensive experiments and analyses on ImageNet, which show that the proposed loss effectively improves the generalizability of trustworthiness predictors. The code and pre-trained trustworthiness predictors for reproducibility are available at https://github.com/luoyan407/predict_trustworthiness .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d46acfb3-4d21-4e74-b2d1-c6f53357c4b3Cited by top-tier papers8
- Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust ClassifiersJonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso et al.NeurIPS 2024 · 12 citations
- RCL: Reliable Continual Learning for Unified Failure DetectionFei Zhu, Zhen Cheng, Xu-Yao Zhang, Cheng-Lin Liu et al.CVPR 2024 · 5 citations
- Verify your Labels! Trustworthy Predictions and Datasets via Confidence ScoresTorsten Krauß, Jasper Stang, Alexandra DmitrienkoUSENIX Security 2024 · 3 citations
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 3 citations
- Adaptive Confidence Regularization for Multimodal Failure DetectionMoru Liu, Hao Dong, Olga Fink, Mario TrappCVPR 2026 · 1 citation
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Dynamic Curriculum Learning for Imbalanced Data ClassificationYiru Wang, Weihao Gan, Jie Yang, Wei Wu et al.ICCV 2019 · 263 citations
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 184 citations
- Metric-Free Individual Fairness in Online LearningYahav Bechavod, Christopher Jung, Zhiwei Steven WuNeurIPS 2020 · 57 citations
- M2m: Imbalanced Classification via Major-to-Minor TranslationJaehyung Kim, Jongheon Jeong, Jinwoo ShinCVPR 2020
Related papers
- Anchor Loss: Modulating Loss Scale Based on Prediction DifficultySerim Ryou, Seong-Gyun Jeong, Pietro PeronaICCV 2019 · 46 citations
- Generative Classifiers as a Basis for Trustworthy Image ClassificationRadek Mackowiak, Lynton Ardizzone, Ullrich Köthe, Carsten RotherCVPR 2021
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Improving Adversarial Robustness via Probabilistically Compact Loss with Logit ConstraintsXin Li, Xiangrui Li, Deng Pan, Dongxiao ZhuAAAI 2021 · 17 citations
- Attack-inspired Calibration Loss for Calibrating Crack RecognitionZhuangzhuang Chen, Qiangyu Chen, Jiahao Zhang, Zhiliang Lin et al.AAAI 2025 · 2 citations
