MediSee: Reasoning-Based Pixel-Level Perception in Medical Images
Qinyue Tong, Ziqian Lu, Jun Liu, Yangming Zheng, Zhe-Ming Lu
Abstract
Despite progress in pixel-level medical image perception, existing methods remain task-specific or depend on precise prompts like bounding boxes or text. However, the need for medical knowledge limits accessibility for the general public, who are more likely to use logically reasoned oral queries than domain-specific inputs. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for MedSD. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods. The MediSee project can be found here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and SegmentationYankai Jiang, Qiaoru Li, Binlu Xu, Haoran Sun et al.CVPR 2026 · 9 citations
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity MasksLingran Song, Yucheng Zhou, Jianbing ShenAAAI 2026
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
Related papers
- Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelAndong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang et al.CVPR 2025
- MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputYanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan et al.CVPR 2025
- MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning SegmentationDonggon Jang, Yucheol Cho, Suin Lee, Taehyeon Kim et al.ICLR 2025
- LISA: Reasoning Segmentation via Large Language ModelXin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li et al.CVPR 2024
- Refer to Any Segmentation Mask Group with Vision-Language PromptsShengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu et al.ICCV 2025
