Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation System
Yuyoung Kim, Minjung Kim, Saebyeol Kim, Sooyoun Cho, Jinwoo Kim
Abstract
Health artificial intelligence (AI) is often developed in high-stakes, data-scarce contexts, where both clinical validity and patient comprehension are critical; however, rigorous, multi-level evaluation of explanations in real-world patient-facing settings remains challenging. To enhance patient understanding and trust, we propose a practical blueprint for designing and evaluating medically aligned, patient-centered explanation (MAP-X). We propose this blueprint through MAP-X, a system that employs a large language model (LLM) with retrieval-augmented generation (RAG) to translate clinical assessments into an understandable interface. We conducted a three-phase evaluation following a multi-level validation framework: a functional evaluation of faithfulness, a clinician evaluation of workflow suitability, and a patient evaluation of perceived understanding and trust. Our findings suggest that MAP-X may support clinical adoption. In the patient study, MAP-X showed higher reported trust and a positive trend in explanation satisfaction. Interviews suggested clearer understanding of assessment results. Overall, MAP-X produced clinically relevant explanations with reasonable faithfulness and usability. Clinician oversight remains necessary.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e9675f02-e4a4-42ae-93f3-70addb73729fRelated papers
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz et al.EMNLP 2025 · 2 citations
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu et al.EMNLP 2025 · 1 citation
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung et al.ACL 2026 · 9 citations
- Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewZhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin et al.ISSTA 2026
- Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case PreparationJiachen Li, Amanda K. Hall, Ruican Rachel Zhong, Selin S. Everett et al.CHI 2026 · 1 citation
