AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
Charles Nimo, Tobi Olatunji, Abraham Toluwase Owodunni, Tassallah Abdullahi, Emmanuel Ayodele, Mardhiyah Sanni, Ezinwanne C. Aka, Folafunmi Omofoye, Foutse Yuehgoh, Timothy Faniran, Bonaventure F. P. Dossou, Moshood O. Yekini
摘要
Recent advancements in large language model(LLM) performance on medical multiple choice question (MCQ) benchmarks have stimulated interest from healthcare providers and patients globally. Particularly in low-and middle-income countries (LMICs) facing acute physician shortages and lack of specialists, LLMs offer a potentially scalable pathway to enhance healthcare access and reduce costs. However, their effectiveness in the Global South, especially across the African continent, remains to be established. In this work, we introduce AfriMed-QA, the first large scale Pan-African English multi-specialty medical Question-Answering (QA) dataset, 15,000 questions (open and closed-ended) sourced from over 60 medical schools across 16 countries, covering 32 medical specialties. We further evaluate 30 LLMs across multiple axes including correctness and demographic bias. Our findings show significant performance variation across specialties and geographies, MCQ performance clearly lags USMLE (MedQA). We find that biomedical LLMs underperform general models and smaller edge-friendly LLMs struggle to achieve a passing score. Interestingly, human evaluations show a consistent consumer preference for LLM answers and explanations when compared with clinician answers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical ReasoningXilin Dang, Kexin Chen, Xiaorui Su, Ayush Noori 等ICLR 2026 · 被引用 6 次
- MedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad 等ICLR 2026 · 被引用 4 次
- KGARevion: An AI Agent for Knowledge-Intensive Biomedical QAXiaorui Su, Yibo Wang, Shanghua Gao, Xiaolong Liu 等ICLR 2025 · 被引用 4 次
- AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMsXiang Feng, Wentao Jiang, Zengmao Wang, Yong Luo 等ICLR 2026 · 被引用 2 次
- Multimodal Medical Code TokenizerXiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson 等ICML 2025 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Africa Health Check: Probing Cultural Bias in Medical LLMsCharles Nimo, Shuheng Liu, Irfan Essa, Michael L. BestEMNLP 2025
- Afri-MCQA: Multimodal Cultural Question Answering for African LanguagesAtnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime 等ACL 2026 · 被引用 2 次
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao 等CVPR 2024
- MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering DatasetJing Li, Shangping Zhong, Kaizhi ChenEMNLP 2021 · 被引用 24 次
- MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsMohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi SoltanolkotabiICLR 2025
