CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
Ekaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio, Serena Villata, Rodrigo Agerri
Abstract
Explaining Artificial Intelligence (AI) decisions is a major challenge nowadays in AI, in particular when applied to sensitive scenarios like medicine and law. However, the need to explain the rationale behind decisions is a main issue also for human-based deliberation as it is important to justify why a certain decision has been taken. Resident medical doctors for instance are required not only to provide a (possibly correct) diagnosis, but also to explain how they reached a certain conclusion. Developing new tools to aid residents to train their explanation skills is therefore a central objective of AI in education. In this paper, we follow this direction, and we present, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors. These explanations have been manually annotated with argument components (i.e., premise, claim) and argument relations (i.e., attack, support). The Multilingual CasiMedicosarg dataset consists of 558 clinical cases in four languages (English, Spanish, French, Italian) with explanations, where we annotated 5021 claims, 2313 premises, 2431 support relations, and 1106 attack relations. We conclude by showing how competitive baselines perform over this challenging dataset for the argument mining task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 15 citations
- Argument Mining in Data Scarce Settings: Cross-lingual Transfer and Few-shot TechniquesAnar Yeginbergen, Maite Oronoz, Rodrigo AgerriACL 2024
Related papers
- e-CARE: a New Dataset for Exploring Explainable Causal ReasoningLi Du, Xiao Ding, Kai Xiong, Ting Liu et al.ACL 2022
- MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsVishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu et al.EMNLP 2024 · 1 citation
- MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-ReasoningFan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang et al.ACL 2026 · 1 citation
- Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QAJustice Ou, Tinglin Huang, Yilun Zhao, Ziyang Yu et al.ACL 2026 · 9 citations
- CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical ReasoningEric Onyame, Akash Ghosh, Subhadip Baidya, Sriparna Saha et al.ACL 2026 · 7 citations
