CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
Chinedu Innocent Nwoye, Kareem Elgohary, Anvita Srinivas, Fauzan Zaid, Joël L. Lavanchy, Nicolas Padoy
Abstract
Tool tracking in surgical videos is essential for advancing computer-assisted interventions, such as skill assessment, safety zone estimation, and human-machine collaboration. However, the lack of context-rich datasets limits AI applications in this field. Existing datasets rely on overly generic tracking formalizations that fail to capture surgical-specific dynamics, such as tools moving out of the camera's view or exiting the body. This results in less clinically relevant trajectories and a lack of flexibility for real-world surgical applications. Methods trained on these datasets often struggle with visual challenges such as smoke, reflection, and bleeding, further exposing the limitations of current approaches. We introduce CholecTrack20, a specialized dataset for multi-class, multi-tool tracking in surgical procedures. It redefines tracking formalization with three perspectives: (1) intraoperative, (2) intracorporeal, and (3) visibility, enabling adaptable and clinically meaningful tool trajectories. The dataset comprises 20 full-length surgical videos, annotated at 1 fps, yielding over 35K frames and 65K labeled tool instances. Annotations include spatial location, category, identity, operator, phase, and visual challenges. Benchmarking state-of-the-art methods on Cholec-Track20 reveals significant performance gaps, with current approaches (< 45% HOTA) failing to meet the accuracy required for clinical translation. These findings motivate the need for advanced and intuitive tracking algorithms and establish CholecTrack20 as a foundation for developing robust AI-driven surgical assistance systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question AnsweringYanjun Li, Yuqian Fu, Tianwen Qian, Qi'ao Xu et al.AAAI 2026 · 14 citations
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video UnderstandingYuhao Su, Anwesa Choudhuri, Zhongpai Gao, Benjamin Planche et al.CVPR 2026 · 10 citations
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised DetectionYiliang Chen, Zhixi Li, Cheng Xu, Alex Qinyang Liu et al.ICLR 2026 · 4 citations
- Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionMeng Wei, Kun Yuan, Shi Li, Yue Zhou et al.AAAI 2026 · 1 citation
Builds on5
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan et al.CVPR 2022 · 305 citations
- SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object TrackingYu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang et al.AAAI 2024 · 96 citations
Related papers
- Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkRulin Zhou, Wenlong He, An Wang, Jianhang Zhang et al.AAAI 2026
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical TrainingLe Ma, Thiago Freitas dos Santos, Nadia Magnenat-Thalmann, Katarzyna WacCVPR 2026
- Towards Unified Surgical Skill AssessmentDaochang Liu, Qiyue Li, Tingting Jiang, Yizhou Wang et al.CVPR 2021
- COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue FusionZekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou et al.ICCV 2025 · 1 citation
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion ModelDanush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal, Omid MohareriICML 2026 · 1 citation
