Text Promptable Surgical Instrument Segmentation with Vision-Language Models
Zijian Zhou, Oluwatosin Alabi, Meng Wei, Tom Vercauteren, Miaojing Shi
Abstract
In this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehension of surgical instruments and adaptability to new instrument types. Inspired by recent advancements in vision-language models, we leverage pretrained image and text encoders as our model backbone and design a text promptable mask decoder consisting of attention-and convolution-based prompting schemes for surgical instrument segmentation prediction. Our model leverages multiple text prompts for each surgical instrument through a new mixture of prompts mechanism, resulting in enhanced segmentation performance. Additionally, we introduce a hard instrument area reinforcement module to improve image feature comprehension and segmentation precision. Extensive experiments on several surgical instrument segmentation datasets demonstrate our model's superior performance and promising generalization capability. To our knowledge, this is the first implementation of a promptable approach to surgical instrument segmentation, offering significant potential for practical application in the field of robotic-assisted surgery. Code is available at https://github.com/franciszzj/TP-SIS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2f4c8b1-fba4-4f60-b0b4-b086eb2f7f4cCited by top-tier papers5
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 12 citations
- FlanS: A Foundation Model for Free-Form Language-based Segmentation in Medical ImagesLongchao Da, Rui Wang, Xiaojian Xu, Parminder Bhatia et al.KDD 2025 · 2 citations
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingMing Hu, Kun Yuan, Yaling Shen, Feilong Tang et al.ICCV 2025 · 1 citation
- Language-Guided Salient Object RankingFang Liu, Yuhao Liu, Ke Xu, Shuquan Ye et al.CVPR 2025
- MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical EnvironmentsEge Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram et al.CVPR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- SurgicalSAM: Efficient Class Promptable Surgical Instrument SegmentationWenxi Yue, Jing Zhang, Kun Hu, Yong Xia et al.AAAI 2024 · 142 citations
- Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt FrameworkYu Zhu, Kang Li, Zheng Li, Pheng-Ann HengCVPR 2026 · 1 citation
- Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosNan Xi, Jingjing Meng, Junsong YuanACM MM 2023 · 11 citations
- Domain-Specific Interactive Prompting for Generalized Nuclei ClassificationBinbin Zheng, Aiqiu Wu, Kai Fan, Ao Li et al.ACM MM 2025
- Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionMeng Wei, Kun Yuan, Shi Li, Yue Zhou et al.AAAI 2026 · 1 citation
