Expert Discussions Improve Comprehension of Difficult Cases in Medical Image Assessment
Mike Schaekermann, Carrie J. Cai, Abigail E. Huang, Rory Sayres
Abstract
Medical data labeling workflows critically depend on accurate assessments from human experts. Yet human assessments can vary markedly, even among medical experts. Prior research has demonstrated benefits of labeler training on performance.
Here we utilized two types of labeler training feedback: highlighting incorrect labels for difficult cases ("individual performance" feedback), and expert discussions from adjudication of these cases. We presented ten generalist eye care professionals with either individual performance alone, or individual performance and expert discussions from specialists. Compared to performance feedback alone, seeing expert discussions significantly improved generalists' understanding of the rationale behind the correct diagnosis while motivating changes in their own labeling approach; and also significantly improved average accuracy on one of four pathologies in a held-out test set. This work suggests that image adjudication may provide benefits beyond developing trusted consensus labels, and that exposure to specialist discussions can be an effective training intervention for medical diagnosis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93e469b7-16ee-4d03-b21d-c225fc3d56e2Cited by top-tier papers3
- Lessons Learned from Designing an AI-Enabled Diagnosis Tool for PathologistsHongyan Gu, Jingbin Huang, Lauren Hung, Xiang 'Anthony' ChenCSCW 2021 · 56 citations
- A hunt for the Snark: Annotator Diversity in Data PracticesShivani Kapania, Alex S. Taylor, Ding WangCHI 2023 · 49 citations
- Ambiguity-aware AI Assistants for Medical Data AnalysisMike Schaekermann, Graeme Beaton, Elaheh Sanoubari, Andrew Lim et al.CHI 2020 · 45 citations
Builds on1
Related papers
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- "Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical ReasoningYuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu et al.CHI 2026 · 1 citation
- MEDebiaser: A Human-AI Feedback System for Mitigating Bias in Multi-label Medical Image ClassificationShaohan Shi, Yuheng Shao, Haoran Jiang, Yunjie Yao et al.UIST 2025
- Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingWei Ji, Shuang Yu, Junde Wu, Kai Ma et al.CVPR 2021
- X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic DiagnosisGui Wang, Zehao Zhong, YongSong Zhou, Yudong Li et al.CVPR 2026
