CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric
Lakshmikar R., Ming Ma
Abstract
Accurate and interpretable medical image segmentation remains a major challenge, as existing deep learning models primarily optimize pixel-level accuracy while overlooking positional reasoning—an essential component for automated report generation and clinical interpretability. We introduce CG-Reasoner, a novel centroid-guided cross-modal framework that jointly performs medical image segmentation and positional reasoning. CG-Reasoner integrates a multimodal large language model (LLM), a newly designed light-weight encoder–decoder architecture, and a Text2Centroid module that predicts lesion centroids from reasoning embeddings—enabling the model to produce both accurate segmentation masks and spatially coherent, clinically meaningful reasoning explanations. Furthermore, we propose PRScore (Positional-Reasoning Score), a robust evaluation metric that jointly measures the spatial and semantic alignment between generated reasoning text and segmentation masks. Experiments on six medical datasets across different imaging modalities demonstrate that CG-Reasoner achieves state-of-the-art performance, offering precise segmentation, spatially coherent reasoning, and clinically interpretable visual-textual explanations within a unified framework. The source code is available at https://github.com/lpmm2025/CG-Reasoner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- TinySAM: Pushing the Envelope for Efficient Segment Anything ModelHan Shu, Wenshuo Li, Yehui Tang, Yiman Zhang et al.AAAI 2025 · 57 citations
- PixelLM: Pixel Reasoning with Large Multimodal ModelZhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao et al.CVPR 2024 · 48 citations
- Towards Injecting Medical Visual Knowledge into Multimodal LLMs at ScaleJunying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao et al.EMNLP 2024 · 43 citations
Related papers
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- SegLLM: Multi-round Reasoning Segmentation with Large Language ModelsXudong Wang, Shaolun Zhang, Shufan Li, Kehan Li et al.ICLR 2025
- Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language ModelsShuai Niu, Jing Ma, Hongzhan Lin, Liang Bai et al.ACL 2025 · 5 citations
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtYi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li et al.ACL 2025 · 15 citations
- MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language ModelsAofei Chang, Le Huang, Alex Boyd, parminder bhatia et al.ICML 2026
