MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Fan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang, Junfeng Jiang, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen, Edison Marrese-Taylor, Kazuma Kobayashi
Abstract
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce MED-COREASONER 1 , a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practicegrounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X 2 , a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that MED-COREASONER improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in lowresource languages. Moreover, model distillation and expert evaluation analysis further confirm that MED-COREASONER produces clinically sound and culturally grounded reasoning traces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92dcede5-10ba-462f-b71b-acd09efe74e9Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Chain-of-Dictionary Prompting Elicits Translation in Large Language ModelsHongyuan Lu, Haoran Yang, Haoyang Huang, Dongdong Zhang et al.EMNLP 2024 · 9 citations
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingYuxin Zuo, Shang Qu, Yifei Li, Zhang-Ren Chen et al.ICML 2025
Related papers
- MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng et al.EMNLP 2025 · 4 citations
- CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical ReasoningEric Onyame, Akash Ghosh, Subhadip Baidya, Sriparna Saha et al.ACL 2026 · 7 citations
- Eliciting Better Multilingual Structured Reasoning from LLMs through CodeBryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas et al.ACL 2024
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 1 citation
- From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlasZhaokun Yan, Shan Xu, Wuzheng Dong, Zhaohan Liu et al.ICML 2026
