TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
Fang Li, Shihao Zou, Weixin Si, Yang Gao, Shuai Li, Aimin Hao
摘要
Understanding complex surgical scenes requires recognizing multiple interdependent entities—such as instruments, actions, and targets—and maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and inter-frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class-specific spatial priors are first extracted through a multi-scale encoder. Then, these priors are refined by a Label Correlation Modeling module with multi-scale class activation map-guided relational extraction (MS-CAMRE), enabling the model to capture both static co-occurrence and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal–Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model’s capability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProStaTD datasets show that our method achieves state-of-the-art (SOTA) performance, improving by 5.1% and 7.8%, respectively. Moreover, on the TCER metric, our approach yields over 36% and 25% relative reductions on the two datasets, respectively, underscoring the effectiveness of our framework in temporal–relational co-reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam 等CVPR 2022 · 被引用 699 次
- UniFormerV2: Unlocking the Potential of Image ViTs for Video UnderstandingKunchang Li, Yali Wang, Yinan He, Yizhuo Li 等ICCV 2023 · 被引用 85 次
- Scene-Aware Label Graph Learning for Multi-Label Image ClassificationXuelin Zhu, Jian Liu, Weijia Liu, Jiawei Ge 等ICCV 2023 · 被引用 39 次
- Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosNan Xi, Jingjing Meng, Junsong YuanACM MM 2023 · 被引用 11 次
相关 Paper
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised DetectionYiliang Chen, Zhixi Li, Cheng Xu, Alex Qinyang Liu 等ICLR 2026 · 被引用 4 次
- Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer TrackingKun Hu, Wenjing Yang, Wanrong Huang, Xianchen Zhou 等AAAI 2024 · 被引用 15 次
- A Speaker-Aware Co-Attention Framework for Medical Dialogue Information ExtractionYuan Xia, Zhenhui Shi, Jingbo Zhou, Jiayu Xu 等EMNLP 2022 · 被引用 6 次
- A Trigger-Sense Memory Flow Framework for Joint Entity and Relation ExtractionYongliang Shen, Xinyin Ma, Yechun Tang, Weiming LuWWW 2021 · 被引用 72 次
- Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer ApproachRochana Chaturvedi, Peyman Baghershahi, Sourav Medya, Barbara Di EugenioACL 2025
