TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
Fang Li, Shihao Zou, Weixin Si, Yang Gao, Shuai Li, Aimin Hao
Abstract
Understanding complex surgical scenes requires recognizing multiple interdependent entities—such as instruments, actions, and targets—and maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and inter-frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class-specific spatial priors are first extracted through a multi-scale encoder. Then, these priors are refined by a Label Correlation Modeling module with multi-scale class activation map-guided relational extraction (MS-CAMRE), enabling the model to capture both static co-occurrence and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal–Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model’s capability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProStaTD datasets show that our method achieves state-of-the-art (SOTA) performance, improving by 5.1% and 7.8%, respectively. Moreover, on the TCER metric, our approach yields over 36% and 25% relative reductions on the two datasets, respectively, underscoring the effectiveness of our framework in temporal–relational co-reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam et al.CVPR 2022 · 699 citations
- UniFormerV2: Unlocking the Potential of Image ViTs for Video UnderstandingKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.ICCV 2023 · 85 citations
- Scene-Aware Label Graph Learning for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Weijia Liu, Jiawei Ge et al.ICCV 2023 · 39 citations
- Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosNan Xi, Jingjing Meng, Junsong YuanACM MM 2023 · 11 citations
Related papers
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised DetectionYiliang Chen, Zhixi Li, Cheng Xu, Alex Qinyang Liu et al.ICLR 2026 · 4 citations
- Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer TrackingKun Hu, Wenjing Yang, Wanrong Huang, Xianchen Zhou et al.AAAI 2024 · 15 citations
- A Speaker-Aware Co-Attention Framework for Medical Dialogue Information ExtractionYuan Xia, Zhenhui Shi, Jingbo Zhou, Jiayu Xu et al.EMNLP 2022 · 6 citations
- A Trigger-Sense Memory Flow Framework for Joint Entity and Relation ExtractionYongliang Shen, Xinyin Ma, Yechun Tang, Weiming LuWWW 2021 · 72 citations
- Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer ApproachRochana Chaturvedi, Peyman Baghershahi, Sourav Medya, Barbara Di EugenioACL 2025
