GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI
Tianbin Li, Yanzhou Su, Wei Li, Bin Fu, Zhe Chen, Ziyan Huang, Guoan Wang, Chenglong Ma, Ying Chen, Ming Hu, Yanjun Li, Pengcheng Chen
Abstract
Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis. Recent advancements in Large-scale Vision-Language Models (LVLMs) have driven progress in image recognition, natural language processing, and multimodal tasks, leveraging the power of multimodal datasets. In the medical field (general medical AI, GMAI), as these technologies mature, the need for accurate processing of diverse data-such as medical images, clinical text, and structured records-has become critical for reliable diagnostic and treatment decisions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5b50980-ff31-478e-9e73-a2f0e3342806Cited by top-tier papers5
- MedMO: Grounding and Understanding Multimodal Large Language Model for Medical ImagesAnkan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal et al.CVPR 2026 · 13 citations
- UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-AnalysisJunzhi Ning, Wei Li, Cheng Tang, Jiashi Lin et al.ICML 2026 · 13 citations
- MedEyes: Learning Dynamic Visual Focus for Medical Progressive DiagnosisChunzheng Zhu, Yangfang Lin, Shen Chen, Yijun Wang et al.AAAI 2026 · 8 citations
- MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksWenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan et al.ACM MM 2025 · 5 citations
- Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation DataXun Zhu, Fanbin Mo, Zheng Zhang, Jiaxi Wang et al.ACM MM 2025
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- VILA-M3: Enhancing Vision-Language Models with Medical Expert KnowledgeVishwesh Nath, Wenqi Li, Dong Yang, Andriy Myronenko et al.CVPR 2025
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
- Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical TasksLehan Wang, Haonan Wang, Honglong Yang, Jiaji Mao et al.ICLR 2025
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu et al.ICCV 2025 · 10 citations
- MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputYanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan et al.CVPR 2025
