Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
Keer Lu, Zheng Liang, Da Pan, Shusen Zhang, Guosheng Dong, Huang Leng, Bin Cui, Zhonghai Wu, Wentao Zhang
Abstract
Large Language Models (LLMs) have exhibited remarkable capabilities in clinical scenarios. Despite their potential, existing works face challenges when applying LLMs to medical settings. Strategies relying on training with medical datasets are highly cost-intensive and may suffer from outdated training data. Leveraging external knowledge bases is a suitable alternative, yet it faces obstacles such as limited retrieval precision and poor effectiveness in answer extraction. These issues collectively prevent LLMs from demonstrating the expected level of proficiency in mastering medical expertise. To address these challenges, we introduce Med-R2, a novel LLM physician framework that adheres to the Evidence-Based Medicine (EBM) process, efficiently integrating retrieval mechanisms as well as the selection and reasoning processes of evidence, thereby enhancing the problem-solving capabilities of LLMs in healthcare scenarios and fostering a trustworthy LLM physician. Our comprehensive experiments indicate that Med-R2 achieves an improvement of 13.27% over vanilla RAG methods and even a 4.55% enhancement compared to fine-tuning strategies, without incurring additional training costs. Furthermore, we find that our LLaMA3.1-70B + Med-R2 surpasses frontier models, including GPT-4o, Claude3.5-Sonnet and DeepSeek-V3 by 1.05%, 6.14% and 1.91%. Med-R2 effectively enhances the capabilities of LLMs in the medical domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c1d9034-7f56-4ab7-92cc-3ee57b3cc7f6Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
Related papers
- Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process RewardsJaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim et al.EMNLP 2025
- MIRA: A Novel Framework for Fusing Modalities in Medical RAGJinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen et al.ACM MM 2025 · 5 citations
- CAMEC: Complexity-Aware Multi-Expert Collaboration for Reliable Chinese Medical Question AnsweringYukang Wu, Xiyuan Jia, Jiayi Wu, Hongchen Yu et al.ACL 2026
- UR² : Unify RAG and Reasoning through Reinforcement LearningWeitao Li, Boran Xiang, Xiaolong Wang, Jingyi Ren et al.ACL 2026 · 1 citation
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question AnsweringLei Li, Xiao Zhou, Yingying Zhang, Xian WuWWW 2026
