MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak
Abstract
achieves an average improvement of +6.6% over Fleming-VL-8B, with gains of +6.0% on MMMU-Med, +9.8% on PMC-VQA, and +21.3% on MedXpertQA. For text-based QA, it attains +14.4% over Fleming-VL-8B, driven by +8.4% on MMLU-Med and +30.1% on MedQA. In medical report generation, MedMO-8B-Next delivers +6.7% on MIMIC-CXR. Moreover, it exhibits strong grounding capability with a Bacteria IoU of 56.1, representing a +47.8 IoU gain over Fleming-VL-8B, underscoring its robust spatial reasoning and localization performance. MedMO-4B-Next remains highly competitive at its smaller scale, surpassing Fleming-VL-8B across VQA, QA, and report generation benchmarks. Evaluations across radiology, ophthalmology, and pathology microscopy confirm MedMO's broad cross-modality generalization.
We further conduct comprehensive experiments and analyses on data curation, training, and alignment strategies, providing a transparent and reproducible framework for future medical MLLM development. Extensive evaluations demonstrate that MedMO achieves state-of-the-art (SOTA) performance across diverse benchmarks, surpassing prior open and proprietary systems on tasks including medical VQA, report generation, and diagnostic reasoning.
Our main contributions are summarized as follows:
• We develop a powerful open-source post-trained multimodal large VLM, MedMO, designed for comprehensive medical image understanding and grounding.
• We curate over 26M multimodal medical and biomedical samples from 45 datasets and establish a multi-stage posttraining that progressively enhances cross-modal alignment and reasoning. This provides a scalable roadmap toward a generalist foundation model for medical.
• To evaluate VLM performance on detection tasks, we construct a dedicated Cell dataset from opensource microscopy images with varying sizes, shapes, and densities.
• We conduct extensive experiments and analyses across data and methodology dimensions, providing an open benchmark for future multimodal medical LLM research and training recipes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 127342fa-2250-4884-b283-760ff2b6d1d3Builds on10
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- FairCLIP: Harnessing Fairness in Vision-Language LearningYan Luo, Min Shi, Muhammad Osama Khan, Muhammad Muneeb Afzal et al.CVPR 2024 · 37 citations
- Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology VideosMehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna et al.CVPR 2024 · 37 citations
- GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AITianbin Li, Yanzhou Su, Wei Li, Bin Fu et al.AAAI 2026 · 1 citation
Related papers
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningHaozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu et al.CVPR 2026 · 15 citations
- Towards a Multimodal Large Language Model with Pixel-Level Insight for BiomedicineXiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang et al.AAAI 2025 · 32 citations
- How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical ImagesGuimeng Liu, Tianze Yu, Somayeh Ebrahimkhani, Lin Zhi Zheng Shawn et al.ICLR 2026 · 3 citations
- Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoEXun Zhu, Ying Hu, Fanbin Mo, Miao Li et al.NeurIPS 2024 · 29 citations
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingDavid Dai, Peilin Chen, Chanakya Ekbote, Paul Pu LiangNeurIPS 2025 · 48 citations
