High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive Assessment
Abhay Sheel Anand, Deepak Ganesan, Ravi Karkar
Abstract
Foundation models tested for clinical practice using human-designed metrics may mask fundamental differences in information processing. We investigated this using the clock drawing test (CDT), a cognitive screening tool. Three foundation models achieved 94% accuracy on conventional metrics, matching experts. However, upon decomposing the CDT into 24 questions across five cognitive domains, results diverged significantly. In cases with unanimous model agreement, they still disagreed with human raters in 22% cases. Performance varied drastically with 88% alignment with humans on rule-based executive questions but only 46% on context-dependent anticipatory thinking questions. We observed that models abstained three times more than humans, primarily owing to poor data quality. These findings show standard clinical evaluation metrics fail to capture how foundation models process information. High aggregate accuracy obscures component-level failures. We contribute a systematic evaluation of frontier models’ healthcare capabilities, demonstrate theory-driven task decomposition, and discuss design implications for better human-AI collaborative systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get de1515e6-4737-4fce-ba53-6767cf0182bcRelated papers
- Beyond the Clock: Exploring Multimodal Behavior Markers of Mild Cognitive Impairment in Older Adults during Clock Drawing TestLingjie Fan, Junhan Zhao, Yongji Wu, Fengyi Wang et al.UbiComp 2026 · 3 citations
- Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in MedicineMaxime Griot, Jean Vanderdonckt, Demet Yüksel, Coralie HemptinneACL 2025
- Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial IntelligenceYanan Wang, Shuaicong Hu, Jian Liu, Guohui Zhou et al.ICML 2026
- An Investigation of Memorization Risk in Healthcare Foundation ModelsSana Tonekaboni, Lena Stempfle, Adibvafa Fallahpour, Walter Gerych et al.NeurIPS 2025 · 7 citations
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
