Text Promptable Surgical Instrument Segmentation with Vision-Language Models
Zijian Zhou, Oluwatosin Alabi, Meng Wei, Tom Vercauteren, Miaojing Shi
摘要
In this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehension of surgical instruments and adaptability to new instrument types. Inspired by recent advancements in vision-language models, we leverage pretrained image and text encoders as our model backbone and design a text promptable mask decoder consisting of attention-and convolution-based prompting schemes for surgical instrument segmentation prediction. Our model leverages multiple text prompts for each surgical instrument through a new mixture of prompts mechanism, resulting in enhanced segmentation performance. Additionally, we introduce a hard instrument area reinforcement module to improve image feature comprehension and segmentation precision. Extensive experiments on several surgical instrument segmentation datasets demonstrate our model's superior performance and promising generalization capability. To our knowledge, this is the first implementation of a promptable approach to surgical instrument segmentation, offering significant potential for practical application in the field of robotic-assisted surgery. Code is available at https://github.com/franciszzj/TP-SIS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 被引用 12 次
- FlanS: A Foundation Model for Free-Form Language-based Segmentation in Medical ImagesLongchao Da, Rui Wang, Xiaojian Xu, Parminder Bhatia 等KDD 2025 · 被引用 2 次
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingMing Hu, Kun Yuan, Yaling Shen, Feilong Tang 等ICCV 2025 · 被引用 1 次
- Language-Guided Salient Object RankingFang Liu, Yuhao Liu, Ke Xu, Shuquan Ye 等CVPR 2025
- MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical EnvironmentsEge Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram 等CVPR 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
相关 Paper
- SurgicalSAM: Efficient Class Promptable Surgical Instrument SegmentationWenxi Yue, Jing Zhang, Kun Hu, Yong Xia 等AAAI 2024 · 被引用 142 次
- Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt FrameworkYu Zhu, Kang Li, Zheng Li, Pheng-Ann HengCVPR 2026 · 被引用 1 次
- Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosNan Xi, Jingjing Meng, Junsong YuanACM MM 2023 · 被引用 11 次
- Domain-Specific Interactive Prompting for Generalized Nuclei ClassificationBinbin Zheng, Aiqiu Wu, Kai Fan, Ao Li 等ACM MM 2025
- Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionMeng Wei, Kun Yuan, Shi Li, Yue Zhou 等AAAI 2026 · 被引用 1 次
