MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Fan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang, Junfeng Jiang, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen, Edison Marrese-Taylor, Kazuma Kobayashi
摘要
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce MED-COREASONER 1 , a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practicegrounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X 2 , a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that MED-COREASONER improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in lowresource languages. Moreover, model distillation and expert evaluation analysis further confirm that MED-COREASONER produces clinically sound and culturally grounded reasoning traces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan 等NeurIPS 2024 · 被引用 291 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Chain-of-Dictionary Prompting Elicits Translation in Large Language ModelsHongyuan Lu, Haoran Yang, Haoyang Huang, Dongdong Zhang 等EMNLP 2024 · 被引用 9 次
- MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingYuxin Zuo, Shang Qu, Yifei Li, Zhang-Ren Chen 等ICML 2025
相关 Paper
- MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng 等EMNLP 2025 · 被引用 4 次
- CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical ReasoningEric Onyame, Akash Ghosh, Subhadip Baidya, Sriparna Saha 等ACL 2026 · 被引用 7 次
- Eliciting Better Multilingual Structured Reasoning from LLMs through CodeBryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas 等ACL 2024
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 被引用 1 次
- From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlasZhaokun Yan, Shan Xu, Wuzheng Dong, Zhaohan Liu 等ICML 2026
