CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction
Sizhe Wang, Ziqi Xu, Claire Najjuuko, Charles Alba, Chenyang Lu
Abstract
Clinical language models (LMs) are increasingly applied to support clinical risk prediction from free-text notes, yet their uncertainty estimates often remain poorly calibrated and clinically unreliable. In this work, we propose Clinical Uncertainty Risk Alignment (CURA), a framework that aligns clinical LM-based risk estimates and uncertainty with both individual error likelihoods and cohort-level ambiguities. CURA first fine-tunes domain-specific clinical LMs to obtain task-adapted patient embeddings, and then performs uncertainty fine-tuning of a multi-head classifier using a bi-level uncertainty objective. Specifically, an individuallevel calibration term aligns predictive uncertainty with each patient's likelihood of error, while a cohort-aware regularizer pulls risk estimates toward event rates in their local neighborhoods in the embedding space and places extra weight on ambiguous cohorts near the decision boundary. We further show that this cohort-aware term can be interpreted as a crossentropy loss with neighborhood-informed soft labels, providing a label-smoothing view of our method. Extensive experiments on MIMIC-IV clinical risk prediction tasks across various clinical LMs show that CURA consistently improves calibration metrics without substantially compromising discrimination. Further analysis illustrates that CURA reduces overconfident false reassurance and yields more trustworthy uncertainty estimates for downstream clinical decision support. Our code is available at: https://github.com/sizhe04/CURA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06de8535-a8ea-4876-af61-024965308b71Builds on5
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins et al.NeurIPS 2024 · 124 citations
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsKaitlyn Zhou, Dan Jurafsky, Tatsunori HashimotoEMNLP 2023 · 29 citations
- Deep Bayesian Active Learning for Preference Modeling in Large Language ModelsLuckeciano Carvalho Melo, Panagiotis Tigas, Alessandro Abate, Yarin GalNeurIPS 2024 · 25 citations
- Taming Overconfidence in LLMs: Reward Calibration in RLHFJixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin HuangICLR 2025
- Reward Uncertainty for Exploration in Preference-based Reinforcement LearningXinran Liang, Katherine Shu, Kimin Lee, Pieter AbbeelICLR 2022
Related papers
- Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language ModelsXiao Liang, Di Wang, Zhicheng Jiao, Ronghan Li et al.ICCV 2025 · 2 citations
- CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report GenerationPablo Messina, Andrés Villa, Juan Leon Alcazar, Karen Sanchez et al.CVPR 2026 · 1 citation
- ViLU: Learning Vision-Language Uncertainties for Failure PredictionMarc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon et al.ICCV 2025 · 1 citation
- CARER - ClinicAl Reasoning-Enhanced Representation for Temporal Health Risk PredictionTuan Nguyen, Thanh Trung Huynh, Minh Hieu Phan, Quoc Viet Hung Nguyen et al.EMNLP 2024 · 1 citation
- MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMsZhan Qu, Michael FärberACL 2026 · 1 citation
