CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation
Zhongzhen Huang, Yankai Jiang, Rongzhao Zhang, Shaoting Zhang, Xiaofan Zhang
摘要
Existing promptable segmentation methods in the medical imaging field primarily consider either textual or visual prompts to segment relevant objects, yet they often fall short when addressing anomalies in medical images, like tumors, which may vary greatly in shape, size, and appearance. Recognizing the complexity of medical scenarios and the limitations of textual or visual prompts, we propose a novel dual-prompt schema that leverages the complementary strengths of visual and textual prompts for segmenting various organs and tumors. Specifically, we introduce CAT, an innovative model that Coordinates Anatomical prompts derived from 3D cropped images with Textual prompts enriched by medical domain knowledge. The model architecture adopts a general query-based design, where prompt queries facilitate segmentation queries for mask prediction. To synergize two types of prompts within a unified framework, we implement a ShareRefiner, which refines both segmentation and prompt queries while disentangling the two types of prompts. Trained on a consortium of 10 public CT datasets, CAT demonstrates superior performance in multiple segmentation tasks. Further validation on a specialized in-house dataset reveals the remarkable capacity of segmenting tumors across multiple cancer stages. This approach confirms that coordinating multimodal prompts is a promising avenue for addressing complex scenarios in the medical domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VoxTell: Free-Text Promptable Universal 3D Medical Image SegmentationMaximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee 等CVPR 2026 · 被引用 22 次
- CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationXinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab 等ACM MM 2025 · 被引用 3 次
- Mitigating Entity Hallucinations in 3D Radiology Report Generation via Dual-Stream AlignmentLingyu Zhou, Yue Yu, Zhang Yi, Xiuyuan XuAAAI 2026
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT ScansHeng Guo, Jianfeng Zhang, Jiaxing Huang, Tony C. W. Mok 等AAAI 2025 · 被引用 12 次
- ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-PromptingYankai Jiang, Zhongzhen Huang, Rongzhao Zhang, Xiaofan Zhang 等CVPR 2024
- K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation ModelBangwei Guo, Yunhe Gao, Meng Ye, Difei Gu 等ICLR 2026 · 被引用 2 次
- GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationLinrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan ZhangAAAI 2026
- Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image SegmentationAishik Konwer, Zhijian Yang, Erhan Bas, Cao Xiao 等CVPR 2025
