MedDialog: Large-scale Medical Dialogue Datasets
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu
Abstract
Medical dialogue systems are promising in assisting in telemedicine to increase access to healthcare services, improve the quality of patient care, and reduce medical costs. To facilitate the research and development of medical dialogue systems, we build large-scale medical dialogue datasets -MedDialog, which contain 1) a Chinese dataset with 3.4 million conversations between patients and doctors, 11.3 million utterances, 660.2 million tokens, covering 172 specialties of diseases, and 2) an English dataset with 0.26 million conversations, 0.51 million utterances, 44.53 million tokens, covering 96 specialties of diseases. To our best knowledge, MedDialog is the largest medical dialogue dataset to date. We pretrain several dialogue generation models on the Chinese MedDialog dataset, including Transformer, GPT, BERT-GPT, and compare their performance. It is shown that models trained on MedDialog are able to generate clinically correct and human-like medical dialogues. We also study the transferability of models trained on MedDialog to lowresource medical dialogue generation tasks. It is shown that via transfer learning which finetunes the models pretrained on MedDialog, the performance on medical dialogue generation tasks with small datasets can be greatly improved, as shown in human evaluation and automatic evaluation. The datasets and code are available at https://github.com/UCSD- AI4H/Medical-Dialogue-System
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5169880b-486e-400e-b7ae-9df3e1fa6257Cited by top-tier papers20
- DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert KnowledgeBufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu et al.UbiComp 2025 · 60 citations
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun et al.SIGIR 2022 · 41 citations
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt et al.ACL 2023 · 19 citations
- Doctor Recommendation in Online Health Forums via Expertise LearningXiaoxin Lu, Yubo Zhang, Jing Li, Shi ZongACL 2022 · 12 citations
- FaMeSumm: Investigating and Improving Faithfulness of Medical SummarizationNan Zhang, Yusen Zhang, Wu Guo, Prasenjit Mitra et al.EMNLP 2023 · 8 citations
Builds on2
- Generative Adversarial Regularized Mutual Information Policy Gradient Framework for Automatic DiagnosisYuan Xia, Jingbo Zhou, Zhenhui Shi, Chao Lu et al.AAAI 2020 · 85 citations
- Importance-Aware Learning for Neural Headline EditingQingyang Wu, Lei Li, Hao Zhou, Ying Zeng et al.AAAI 2020 · 17 citations
Related papers
- CDialog: A Multi-turn Covid-19 Conversation Dataset for Entity-Aware Dialog GenerationDeeksha Varshney, Aizan Zafar, Niranshu Kumar Behra, Asif EkbalEMNLP 2022 · 2 citations
- MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsVishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu et al.EMNLP 2024 · 1 citation
- MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech TranslationKhai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui et al.EMNLP 2025
- Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingJiaxin Wen, Hao Zhou, Jian Guan, Jie Zhou et al.EMNLP 2023 · 2 citations
- MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM AgentsCongchi Yin, Feng Li, Shu Zhang, Zike Wang et al.AAAI 2025 · 18 citations
