MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
Wenting Chen, Guolin Huang, Wenxuan Wang, Zhongrui Zhu
Abstract
Despite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis--relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical cases. Existing benchmarks fail to detect this critical failure mode. We introduce MedEinst, a counterfactual benchmark with 5,383 paired clinical cases across 49 diseases. Each pair contains a control case and a"trap"case with altered discriminative evidence that flips the diagnosis. We measure susceptibility via Bias Trap Rate--probability of misdiagnosing traps despite correctly diagnosing controls. Extensive Evaluation of 17 LLMs shows frontier models achieve high baseline accuracy but severe bias trap rates. Thus, we propose ECR-Agent, aligning LLM reasoning with Evidence-Based Medicine standard via two components: (1) Dynamic Causal Inference (DCI) performs structured reasoning through dual-pathway perception, dynamic causal graph reasoning across three levels (association, intervention, counterfactual), and evidence audit for final diagnosis; (2) Critic-Driven Graph and Memory Evolution (CGME) iteratively refines the system by storing validated reasoning paths in an exemplar base and consolidating disease-specific knowledge into evolving illness graphs. Source code is to be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e671671-45c5-45d8-99c8-7ab16b63776fBuilds on1
Related papers
- MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential DiagnosisDaniel Philip Rose, Chia-Chien Hung, Marco Lepri, Israa Alqassem et al.ACL 2025 · 14 citations
- MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive RegulationYu Zhao, Hao Guan, Yongcheng Jing, Ying Zhang et al.ICML 2026
- Beyond Accuracy: Latent Perturbations for Cognitive-Aware DiagnosisYuting Yan, Yinghao Fu, Wendi Ren, Haozhou Gao et al.ICML 2026
- MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowZiyue Wang, Junde Wu, Linghan Cai, Chang Han Low et al.ICLR 2026 · 84 citations
- MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsSiqi Ma, Jiajie Huang, Fan Zhang, Jinlin Wu et al.AAAI 2026 · 10 citations
