ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
Duy M. H. Nguyen, Nghiem Tuong Diep, Trung Nguyen, Hoang-Bao Le, Tai D. Nguyen, Anh-Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, Daniel Sonntag, James Y. Zou, Mathias Niepert
Abstract
State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLAVA-MED and BIOMEDGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce EXGRA-MED, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMa-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, EXGRA-MED matches LLAVA-MED's performance using just 10% of pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BIOMEDGPT and RADFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. Method VQA-RAD SLAKE PathVQA Overall Open Closed Avg. Open Closed Avg. Open Closed Avg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 708bc4e5-9439-47d2-94d3-ea92d7965e49Cited by top-tier papers3
- UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-AnalysisJunzhi Ning, Wei Li, Cheng Tang, Jiashi Lin et al.ICML 2026 · 13 citations
- MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language ModelsAofei Chang, Le Huang, Alex Boyd, parminder bhatia et al.ICML 2026
- MEDA: Medical-Oriented Activation Editing for Hallucination Mitigation in Medical Large Vision-Language ModelTianbo Wang, Yuqing Ma, Lingyan Meng, Zhange Zhang et al.ICML 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image UnderstandingXuanzhao Dong, Wenhui Zhu, Xiwen Chen, Zhipeng Wang et al.CVPR 2026 · 16 citations
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 18 citations
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation AlignmentXing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan et al.AAAI 2026 · 2 citations
