Automatic Construction of Clinical Scoring Systems with LLM Agents
Silas Ruhrberg Estevez, Chris Chiu, Mihaela van der Schaar
Abstract
Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules. While machine-learning models achieve strong performance, many fail to translate into routine clinical use due to misalignment with workflow constraints such as memorability, auditability, and bedside execution. We argue that this gap arises not from insufficient predictive power, but from optimizing over model classes that are incompatible with guideline deployment. Deployable guidelines often take the form of unit-weighted clinical checklists, formed by thresholding the sum of binary rules, but learning such scores requires searching an exponentially large discrete space of possible rule sets. We introduce AgentScore, which performs semantically guided optimization in this space by using LLMs to propose candidate rules and a deterministic, data-grounded verification-and-selection loop to enforce statistical validity and deployability constraints. Across eight clinical prediction tasks, AgentScore outperforms existing score-generation methods and achieves AUROC comparable to more flexible interpretable models despite operating under stronger structural constraints. On two additional externally validated tasks, AgentScore achieves higher discrimination than established guideline-based scores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- FasterRisk: Fast and Accurate Interpretable Risk ScoresJiachang Liu, Chudi Zhong, Boxuan Li, Margo I. Seltzer et al.NeurIPS 2022 · 45 citations
- Learning Optimal Predictive ChecklistsHaoran Zhang, Quaid Morris, Berk Ustun, Marzyeh GhassemiNeurIPS 2021 · 15 citations
- Decision Tree Induction Through LLMs via Semantically-Aware EvolutionTennison Liu, Nicolas Huynh, Mihaela van der SchaarICLR 2025
Related papers
- GLEAN: Guideline-Grounded Evidence Accumulation for High-Stakes Agent VerificationYichi Zhang, Nabeel Seedat, Yinpeng Dong, Peng Cui et al.ICML 2026 · 3 citations
- Universal Guideline-Driven Image Clustering via a Hybrid LLM AgentWenliang Zhong, Rob Barton, Lucas Goncalves, Kushal Kumar et al.CVPR 2026
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsZonghai Yao, Hong YuACL 2026
- ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent CollaborationZixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng et al.WWW 2025 · 34 citations
