MedDialog: Large-scale Medical Dialogue Datasets
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu
摘要
Medical dialogue systems are promising in assisting in telemedicine to increase access to healthcare services, improve the quality of patient care, and reduce medical costs. To facilitate the research and development of medical dialogue systems, we build large-scale medical dialogue datasets -MedDialog, which contain 1) a Chinese dataset with 3.4 million conversations between patients and doctors, 11.3 million utterances, 660.2 million tokens, covering 172 specialties of diseases, and 2) an English dataset with 0.26 million conversations, 0.51 million utterances, 44.53 million tokens, covering 96 specialties of diseases. To our best knowledge, MedDialog is the largest medical dialogue dataset to date. We pretrain several dialogue generation models on the Chinese MedDialog dataset, including Transformer, GPT, BERT-GPT, and compare their performance. It is shown that models trained on MedDialog are able to generate clinically correct and human-like medical dialogues. We also study the transferability of models trained on MedDialog to lowresource medical dialogue generation tasks. It is shown that via transfer learning which finetunes the models pretrained on MedDialog, the performance on medical dialogue generation tasks with small datasets can be greatly improved, as shown in human evaluation and automatic evaluation. The datasets and code are available at https://github.com/UCSD- AI4H/Medical-Dialogue-System
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert KnowledgeBufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu 等UbiComp 2025 · 被引用 60 次
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt 等ACL 2023 · 被引用 19 次
- Doctor Recommendation in Online Health Forums via Expertise LearningXiaoxin Lu, Yubo Zhang, Jing Li, Shi ZongACL 2022 · 被引用 12 次
- FaMeSumm: Investigating and Improving Faithfulness of Medical SummarizationNan Zhang, Yusen Zhang, Wu Guo, Prasenjit Mitra 等EMNLP 2023 · 被引用 8 次
它引用的顶会 Paper2
相关 Paper
- CDialog: A Multi-turn Covid-19 Conversation Dataset for Entity-Aware Dialog GenerationDeeksha Varshney, Aizan Zafar, Niranshu Kumar Behra, Asif EkbalEMNLP 2022 · 被引用 2 次
- MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsVishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu 等EMNLP 2024 · 被引用 1 次
- MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech TranslationKhai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui 等EMNLP 2025
- Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingJiaxin Wen, Hao Zhou, Jian Guan, Jie Zhou 等EMNLP 2023 · 被引用 2 次
- MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM AgentsCongchi Yin, Feng Li, Shu Zhang, Zike Wang 等AAAI 2025 · 被引用 18 次
