MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models
Aofei Chang, Le Huang, Alex Boyd, parminder bhatia, Taha Kass-Hout, Fenglong Ma, Cao Xiao
Abstract
Medical large vision-language models (Med-LVLMs) have recently achieved remarkable progress in vision–language comprehension and medical image segmentation. However, existing models still struggle to unify these two capabilities, which is essential for achieving clinically reasoning that connects visual findings with semantic interpretation. We present MedSIGHT, a unified framework that equips Med-LVLMs with structured, pixel-level understanding for grounded visual comprehension. MedSIGHT introduces a novel Region Perceiver module that produces region-centric tokens, encoding spatial information directly into representation space of the language model. We further propose a medical region codebook into the LLM vocabulary, allowing the model to generate discrete region codes as symbolic representations of anatomical and pathological regions. These codes are decoded through the Region Perceiver to reconstruct segmentation mask, achieving end-to-end spatial grounding. Lastly, MedSIGHT combines Region Perceiver, Codebook and LLM using our proposed progressive training strategy to gradually aligns these modules stably. Trained on only 72K multimodal instruction pairs, MedSIGHT achieves state-of-the-art performance across diverse imaging modalities on both medical comprehension and segmentation tasks. Code and model are publicly available at GitHub .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- X-SAM: From Segment Anything to Any SegmentationHao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang et al.AAAI 2026 · 16 citations
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren et al.NeurIPS 2025 · 47 citations
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language ModelsCong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng et al.ICCV 2025 · 8 citations
