Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic Videos
Nan Xi, Jingjing Meng, Junsong Yuan
Abstract
Surgical triplet recognition aims to recognize surgical activities as triplets (i.e., ), which provides fine-grained information essential for surgical scene understanding. Existing methods for surgical triplet recognition rely on compositional methods that recognize the instrument, verb, and target simultaneously. In contrast, our method, called chain-of-look prompting, casts the problem of surgical triplet recognition as visual prompt generation from large-scale vision-language (VL) models, and explicitly decomposes the task into a series of video reasoning processes. Chain-of-Look prompting is inspired by: (1) the chain-of-thought prompting in natural language processing, which divides a problem into a sequence of intermediate reasoning steps; (2) the inter-dependency between motion and visual appearance in the human vision system. Since surgical activities are conveyed by the actions of physicians, we regard the verbs as the carrier of semantics in surgical endoscopic videos. Additionally, we utilize the BioMed large language model to calibrate the generated visual prompt features for surgical scenarios. Our approach captures the visual reasoning processes underlying surgical activities and achieves better performance compared to the state-of-the-art methods on the largest surgical triplet recognition dataset, CholecT50. The code is available at https://github.com/southnx/CoLSurgical.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens GroundingYue Fan, Lei Ding, Ching-Chen Kuo, Shan Jiang et al.EMNLP 2024 · 1 citation
- TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet RecognitionFang Li, Shihao Zou, Weixin Si, Yang Gao et al.CVPR 2026
Related papers
- SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought BenchmarkGui Wang, YongSong Zhou, Kaijun Deng, Wooi Ping Cheah et al.CVPR 2026
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingMing Hu, Kun Yuan, Yaling Shen, Feilong Tang et al.ICCV 2025 · 1 citation
- SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language ModelYaoqian Li, Xikai Yang, Dunyuan Xu, Yang Yu et al.AAAI 2026 · 4 citations
- MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool CallingWenjie Li, Yujie Zhang, Haoran Sun, Xingqi He et al.ICML 2026 · 3 citations
- Text Promptable Surgical Instrument Segmentation with Vision-Language ModelsZijian Zhou, Oluwatosin Alabi, Meng Wei, Tom Vercauteren et al.NeurIPS 2023 · 56 citations
