MedAraBench: Large-scale Arabic Medical Question Answering Dataset and Benchmark
Mouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad, Mariam Al-Omari, Jihad Mallat, Khaled Saleh, Nizar Habash, Farah Shamout
摘要
Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders efforts to evaluate and advance the multilingual capabilities of Large Language Models (LLMs). In this paper, we introduce MedAraBench, a large-scale dataset consisting of Arabic multiple-choice question-answer pairs across various medical specialties. We constructed the dataset by manually digitizing a large repository of academic materials created by medical professionals in the Arabic-speaking region. We then conducted extensive preprocessing and split the dataset into training and test sets to support future research efforts in the area. To assess the quality of the data, we adopted two frameworks, namely expert human evaluation and LLM-as-a-judge. Our dataset is diverse and of high quality, spanning 19 specialties and five difficulty levels. For benchmarking purposes, we assessed the performance of sixteen state-of-the-art open-source and proprietary models, such as GPT-5, Gemini 2.0 Flash, and Claude 4-Sonnet. Our findings highlight the need for further domain-specific enhancements. We also explore QLoRA fine-tuning on LLaMa-3.1-8B-instruct to assess our dataset's viability. We release the dataset and evaluation scripts to broaden the diversity of medical data benchmarks, expand the scope of evaluation suites for LLMs, and enhance the multilingual capabilities of models for deployment in clinical settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani 等ACL 2025 · 被引用 144 次
- AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark DatasetCharles Nimo, Tobi Olatunji, Abraham Toluwase Owodunni, Tassallah Abdullahi 等ACL 2025 · 被引用 26 次
- ALLaM: Large Language Models for Arabic and EnglishM. Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi 等ICLR 2025 · 被引用 4 次
相关 Paper
- TounsiBench: Benchmarking Large Language Models for Tunisian ArabicSouha Hassine, Asma Arrak, Marouene Addhoum, Steven R. WilsonEMNLP 2025
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo 等AAAI 2024 · 被引用 42 次
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank 等ICLR 2026 · 被引用 24 次
- MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationZexue He, Yu Wang, An Yan, Yao Liu 等EMNLP 2023 · 被引用 7 次
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 被引用 1 次
