Harnessing Biomedical Literature to Calibrate Clinicians' Trust in AI Decision Support Systems
Qian Yang, Yuexing Hao, Kexin Quan, Stephen Yang, Yiran Zhao, Volodymyr Kuleshov, Fei Wang
Abstract
Clinical decision support tools (DSTs), powered by Artificial Intelligence (AI), promise to improve clinicians’ diagnostic and treatment decision-making. However, no AI model is always correct. DSTs must enable clinicians to validate each AI suggestion, convincing them to take the correct suggestions while rejecting its errors. While prior work often tried to do so by explaining AI’s inner workings or performance, we chose a different approach: We investigated how clinicians validated each other’s suggestions in practice (often by referencing scientific literature) and designed a new DST that embraces these naturalistic interactions. This design uses GPT-3 to draw literature evidence that shows the AI suggestions’ robustness and applicability (or the lack thereof). A prototyping study with clinicians from three disease areas proved this approach promising. Clinicians’ interactions with the prototype also revealed new design and research opportunities around (1) harnessing the complementary strengths of literature-based and predictive decision supports; (2) mitigating risks of de-skilling clinicians; and (3) offering low-data decision support with literature.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a1dbd496-d479-450b-a216-964ee7f4a56aCited by top-tier papers23
- Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisShao Zhang, Jianing Yu, Xuhai Xu, Changchang Yin et al.CHI 2024 · 95 citations
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa et al.CHI 2024 · 81 citations
- The Impact of Imperfect XAI on Human-AI Decision-MakingKatelyn Morrison, Philipp Spitzer, Violet Turri, Michelle Feng et al.CSCW 2024 · 60 citations
- "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision MakingShuai Ma, Xinru Wang, Ying Lei, Chuhan Shi et al.CHI 2024 · 54 citations
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
Related papers
- Prompting, Oversight, and Adoption: Physicians' Use of Large Language Models for Diagnostic Reasoning in an LMICUshna Malik, Laiba Intizar Ahmad, Amna Hassan, Izzah Shafique et al.CHI 2026 · 1 citation
- Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical LensMaia L. Jacobs, Jeffrey He, Melanie F. Pradier, Barbara D. Lam et al.CHI 2021 · 171 citations
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu et al.EMNLP 2025 · 1 citation
- Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial IntelligenceYanan Wang, Shuaicong Hu, Jian Liu, Guohui Zhou et al.ICML 2026
- Predicting Clinical Trial Results by Implicit Evidence IntegrationQiao Jin, Chuanqi Tan, Mosha Chen, Xiaozhong Liu et al.EMNLP 2020 · 6 citations
