Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
Gerardo Flores, Alyssa H. Smith, Julia Fukuyama, Ashia C. Wilson
Abstract
Machine learning-based decision support systems are increasingly deployed in clinical settings, where probabilistic scoring functions are used to inform and prioritize patient management decisions. However, widely used scoring rules, such as accuracy and AUC-ROC, fail to adequately reflect key clinical priorities, including calibration, robustness to distributional shifts, and sensitivity to asymmetric error costs. In this work, we propose a principled yet practical evaluation framework for selecting calibrated thresholded classifiers that explicitly accounts for the uncertainty in class prevalences and domain-specific cost asymmetries often found in clinical settings. Building on the theory of proper scoring rules, particularly the Schervish representation, we derive an adjusted variant of cross-entropy (log score) that averages cost-weighted performance over clinically relevant ranges of class balance. The resulting evaluation is simple to apply, sensitive to clinical deployment conditions, and designed to prioritize models that are both calibrated and robust to real-world variations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f0d62f6-61a5-41f7-9b34-410c8d11b06fBuilds on4
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- A Closer Look at AUROC and AUPRC under Class ImbalanceMatthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti et al.NeurIPS 2024 · 191 citations
- A Unified View of Label Shift EstimationSaurabh Garg, Yifan Wu, Sivaraman Balakrishnan, Zachary C. LiptonNeurIPS 2020 · 186 citations
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 22 citations
Related papers
- Principled Algorithms for Optimizing Generalized Metrics in Binary ClassificationAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Reliable Decisions with Threshold CalibrationRoshni Sahoo, Shengjia Zhao, Alyssa Chen, Stefano ErmonNeurIPS 2021 · 35 citations
- The VOROS: Lifting ROC Curves to 3D to Summarize Unbalanced Classifier PerformanceChristopher Ratigan, Lenore CowenAAAI 2025 · 1 citation
- Instance-Level Costs for Nuanced Classifier EvaluationKabir Kang, Steve MussmannICML 2026
- Overcoming Common Flaws in the Evaluation of Selective Classification SystemsJeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner et al.NeurIPS 2024 · 44 citations
