MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive Annotations
Vishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu, Mausam
Abstract
Medical task-oriented dialogue systems can assist doctors by collecting patient medical history, aiding in diagnosis, or guiding treatment selection, thereby reducing doctor burnout and expanding access to medical services. However, doctor-patient dialogue datasets are not readily available, primarily due to privacy regulations. Moreover, existing datasets lack comprehensive annotations involving medical slots and their different attributes, such as symptoms and their onset, progression, and severity. These comprehensive annotations are crucial for accurate diagnosis. Finally, most existing datasets are non-English, limiting their utility for the larger research community. In response, we introduce MediTOD, a new dataset of doctor-patient dialogues in English for the medical history-taking task. Collaborating with doctors, we devise a questionnairebased labeling scheme tailored to the medical domain. Then, medical professionals create the dataset with high-quality comprehensive annotations, capturing medical slots and their attributes. We establish benchmarks in supervised and few-shot settings on MediTOD for natural language understanding, policy learning, and natural language generation subtasks, evaluating models from both TOD and biomedical domains. We release MediTOD resources for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91ef4938-7901-49b5-a3ef-6dba557933ebCited by top-tier papers2
- "Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert CollaborationsYang Wu, Jinhong Yu, Jingwei Xiong, Zhimin Tao et al.ACL 2026 · 1 citation
- Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical NotesYang Zhou, Zhenting Sheng, Mingrui Tan, Yuting Song et al.AAAI 2026
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
- Dialogue State Tracking with a Language Model using Schema-Driven PromptingChia-Hsuan Lee, Hao Cheng, Mari OstendorfEMNLP 2021 · 87 citations
- The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare ApplicationsMina Valizadeh, Natalie PardeACL 2022 · 55 citations
Related papers
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang et al.EMNLP 2020 · 163 citations
- MidMed: Towards Mixed-Type Dialogues for Medical ConsultationXiaoming Shi, Zeming Liu, Chuan Wang, Haitao Leng et al.ACL 2023
- CDialog: A Multi-turn Covid-19 Conversation Dataset for Entity-Aware Dialog GenerationDeeksha Varshney, Aizan Zafar, Niranshu Kumar Behra, Asif EkbalEMNLP 2022 · 2 citations
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 210 citations
- CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative StructuresEkaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio et al.EMNLP 2024 · 2 citations
