DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
Zonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon, Soie Kwon, Hong Yu
摘要
Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education. While recent large language model (LLM) benchmarks emphasize in-visit diagnostic reasoning, they fail to evaluate models' ability to support patients after the visit. We introduce DischargeSim, a novel benchmark that evaluates LLMs on their ability to act as personalized discharge educators. DischargeSim simulates post-visit, multi-turn conversations between LLM-driven DoctorAgents and Pa-tientAgents with diverse psychosocial profiles (e.g., health literacy, education, emotion). Interactions are structured across six clinically grounded discharge topics and assessed along three axes: (1) dialogue quality via automatic and LLM-as-judge evaluation, (2) personalized document generation including free-text summaries and structured AHRQ checklists, and (3) patient comprehension through a downstream multiple-choice exam. Experiments across 18 LLMs reveal significant gaps in discharge education capability, with performance varying widely across patient profiles. Notably, model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization. DischargeSim offers a first step toward benchmarking LLMs in post-visit clinical education and promoting equitable, personalized patient support. 1 . * indicates equal contribution 1 The source code is released at: https://github.com/ michaels6060/DischargeSim with CC-BY-NC 4.0 license.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang 等AAAI 2024 · 被引用 57 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- RARE: Retrieval-Augmented Reasoning Enhancement for Large Language ModelsHieu Tran, Zonghai Yao, Zhichao Yang, Junda Wang 等ACL 2025 · 被引用 27 次
- Unsupervised Commonsense Question Answering with Self-TalkVered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula 等EMNLP 2020 · 被引用 25 次
- StoryER: Automatic Story Evaluation via Ranking, Rating and ReasoningHong Chen, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao 等EMNLP 2022 · 被引用 3 次
相关 Paper
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo 等AAAI 2024 · 被引用 42 次
- ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsYusheng Liao, Shuyang Jiang, Yanfeng Wang, Yu WangACL 2025 · 被引用 14 次
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic EvaluationXiangxu Zhang, Lei Li, Yanyun Zhou, Xiao Zhou 等ACL 2026 · 被引用 3 次
- Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QAJustice Ou, Tinglin Huang, Yilun Zhao, Ziyang Yu 等ACL 2026 · 被引用 9 次
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan 等NeurIPS 2024 · 被引用 291 次
