Human or LLM as Standardized Patients? A Comparative Study in Medical Education
Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Zhou Lei, Qianqian Xie, Benyou Wang
Abstract
Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale. Although large language model (LLM)-based virtual standardized patients (VSPs) have been proposed as an alternative, their behavior remains unstable and lacks rigorous comparison with human standardized patients. We propose EasyMED, a multi-agent VSP framework that separates case-grounded information disclosure from response generation to support stable, inquiryconditioned patient behavior. We also introduce SPBench, a human-grounded benchmark with eight expert-defined criteria for interaction-level evaluation. Experiments show that EasyMED more closely matches human SP behavior than existing VSPs, particularly in case consistency and controlled disclosure. A four-week controlled study further demonstrates learning outcomes comparable to human SP training, with stronger early gains for novice learners and improved flexibility, psychological safety, and cost efficiency. The code is publicly available at https://github.com/ FreedomIntelligence/EasyMED .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e16a21a-059d-4df5-8a15-cf3f6ec550f7Builds on5
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Scaffolding Empathy: Training Counselors with Simulated Patients and Utterance-level Performance VisualizationsIan Steenstra, Farnaz Nouraei, Timothy W. BickmoreCHI 2025 · 30 citations
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling et al.ACL 2024 · 8 citations
- Factual Confidence of LLMs: on Reliability and Robustness of Current EstimatorsMatéo Mahaut, Laura Aina, Paula Czarnowska, Momchil Hardalov et al.ACL 2024 · 7 citations
- LLMs Can Simulate Standardized Patients via Agent CoevolutionZhuoyun Du, Lujie Zheng, Renjun Hu, Yuyang Xu et al.ACL 2025
Related papers
- 'Poker with Play Money': Exploring Psychotherapist Training with Virtual PatientsCynthia M. Baseman, Masum Hasan, Nathaniel Swinger, Sheila A. M. Rauch et al.CSCW 2025 · 2 citations
- ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsYusheng Liao, Shuyang Jiang, Yanfeng Wang, Yu WangACL 2025 · 14 citations
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo et al.EMNLP 2024 · 6 citations
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu et al.EMNLP 2025 · 1 citation
- 3MDBench: Medical Multimodal Multi-agent Dialogue BenchmarkIvan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova et al.EMNLP 2025
