Lune

CVPR2026Top-tier venue

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak

2026Year
13Citations

Abstract

achieves an average improvement of +6.6% over Fleming-VL-8B, with gains of +6.0% on MMMU-Med, +9.8% on PMC-VQA, and +21.3% on MedXpertQA. For text-based QA, it attains +14.4% over Fleming-VL-8B, driven by +8.4% on MMLU-Med and +30.1% on MedQA. In medical report generation, MedMO-8B-Next delivers +6.7% on MIMIC-CXR. Moreover, it exhibits strong grounding capability with a Bacteria IoU of 56.1, representing a +47.8 IoU gain over Fleming-VL-8B, underscoring its robust spatial reasoning and localization performance. MedMO-4B-Next remains highly competitive at its smaller scale, surpassing Fleming-VL-8B across VQA, QA, and report generation benchmarks. Evaluations across radiology, ophthalmology, and pathology microscopy confirm MedMO's broad cross-modality generalization.

We further conduct comprehensive experiments and analyses on data curation, training, and alignment strategies, providing a transparent and reproducible framework for future medical MLLM development. Extensive evaluations demonstrate that MedMO achieves state-of-the-art (SOTA) performance across diverse benchmarks, surpassing prior open and proprietary systems on tasks including medical VQA, report generation, and diagnostic reasoning.

Our main contributions are summarized as follows:

• We develop a powerful open-source post-trained multimodal large VLM, MedMO, designed for comprehensive medical image understanding and grounding.

• We curate over 26M multimodal medical and biomedical samples from 45 datasets and establish a multi-stage posttraining that progressively enhances cross-modal alignment and reasoning. This provides a scalable roadmap toward a generalist foundation model for medical.

• To evaluate VLM performance on detection tasks, we construct a dedicated Cell dataset from opensource microscopy images with varying sizes, shapes, and densities.

• We conduct extensive experiments and analyses across data and methodology dimensions, providing an open benchmark for future multimodal medical LLM research and training recipes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 127342fa-2250-4884-b283-760ff2b6d1d3

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines