Expert load matters: operating networks at high accuracy and low manual effort
Sara Sangalli, Ertunc Erdil, Ender Konukoglu
Abstract
In human-AI collaboration systems for critical applications, in order to ensure minimal error, users should set an operating point based on model confidence to determine when the decision should be delegated to human experts. Samples for which model confidence is lower than the operating point would be manually analysed by experts to avoid mistakes. Such systems can become truly useful only if they consider two aspects: models should be confident only for samples for which they are accurate, and the number of samples delegated to experts should be minimized. The latter aspect is especially crucial for applications where available expert time is limited and expensive, such as healthcare. The trade-off between the model accuracy and the number of samples delegated to experts can be represented by a curve that is similar to an ROC curve, which we refer to as confidence operating characteristic (COC) curve. In this paper, we argue that deep neural networks should be trained by taking into account both accuracy and expert load and, to that end, propose a new complementary loss function for classification that maximizes the area under this COC curve. This promotes simultaneously the increase in network accuracy and the reduction in number of samples delegated to humans. We perform experiments on multiple computer vision and medical image datasets for classification. Our results demonstrate that the proposed loss improves classification accuracy and delegates less number of decisions to experts, achieves better out-of-distribution samples detection and on par calibration performance compared to existing loss functions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- Confidence-aware Contrastive Learning for Selective ClassificationYu-Chang Wu, Shen-Huan Lyu, Haopu Shang, Xiangyu Wang et al.ICML 2024 · 9 citations
Builds on11
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Scaling Out-of-Distribution Detection for Real-World SettingsDan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou et al.ICML 2022 · 653 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Improving model calibration with accuracy versus uncertainty optimizationRanganath Krishnan, Omesh TickooNeurIPS 2020 · 217 citations
Related papers
- Constrained Optimization to Train Neural Networks on Critical and Under-Represented ClassesSara Sangalli, Ertunc Erdil, Andreas M. Hötker, Olivio Donati et al.NeurIPS 2021 · 37 citations
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 184 citations
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 84 citations
- Coverage-Constrained Human-AI Cooperation with Multiple ExpertsZheng Zhang, Cuong C. Nguyen, Kevin Wells, Thanh-Toan Do et al.AAAI 2026
- Probabilistic Learning to Defer: Handling Missing Expert Annotations and Controlling Workload DistributionCuong C. Nguyen, Thanh-Toan Do, Gustavo CarneiroICLR 2025
