GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
Yifan Jiao, Yunhao Li, Junhua Ding, Qing Yang, Song Fu, Heng Fan, Libo Zhang
Abstract
In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple modalities, including the point cloud (PC), RGB image, and depth. This allows GSOT3D to support various 3D tracking tasks, such as single-modal 3D SOT on PC and multi-modal 3D SOT on RGB-PC or RGB-D, and thus greatly broadens research directions for 3D object tracking. To provide highquality per-frame 3D annotations, all sequences are labeled manually with multiple rounds of meticulous inspection and refinement. To our best knowledge, GSOT3D is the largest benchmark dedicated to various generic 3D object tracking tasks. To understand how existing 3D trackers perform and to provide comparisons for future research on GSOT3D, we assess eight representative point cloud-based tracking models. Our evaluation results exhibit that these models heavily degrade on GSOT3D, and more efforts are required for robust and generic 3D object tracking. Besides, to encourage future research, we present a simple yet effective generic 3D tracker, named PROT3D, that localizes the target object via a progressive spatial-temporal network and outperforms all current solutions by a large margin. By releasing GSOT3D, we expect to advance further 3D tracking in future research and applications. Our benchmark and model as well as the evaluation results will be publicly released at our webpage https://github.com/ailovejinx/GSOT3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- PTTR: Relational 3D Point Cloud Object Tracking with TransformerChangqing Zhou, Zhipeng Luo, Yueru Luo, Tianrui Liu et al.CVPR 2022 · 117 citations
- Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsChaoda Zheng, Xu Yan, Jiantao Gao, Weibing Zhao et al.ICCV 2021 · 116 citations
- 3D Siamese Voxel-to-BEV Tracker for Sparse Point CloudsLe Hui, Lingpeng Wang, Mingmei Cheng, Jin Xie et al.NeurIPS 2021 · 105 citations
- Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point CloudsChaoda Zheng, Xu Yan, Haiming Zhang, Baoyuan Wang et al.CVPR 2022 · 100 citations
- Synchronize Feature Extracting and Matching: A Single Branch Framework for 3D Object TrackingTeli Ma, Mengmeng Wang, Jimin Xiao, Huifeng Wu et al.ICCV 2023 · 21 citations
Related papers
- Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object TrackingJie Xiao, Yinchao Ma, Yuyang Tang, Dengqing Yang et al.CVPR 2026
- CDTB: A Color and Depth Visual Object Tracking Dataset and BenchmarkAlan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush et al.ICCV 2019 · 79 citations
- Towards Visual Query Localization in the 3D WorldLiang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li et al.CVPR 2026 · 1 citation
- TrackAny3D: Transferring Pretrained 3D Models for Category-Unified 3D Point Cloud TrackingMengmeng Wang, Haonan Wang, Yulong Li, Xiangjie Kong et al.ICCV 2025 · 2 citations
- PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D TrackingJuntao Liu, Zikun Zhou, Zhuotao Tian, Guangming Lu et al.ICLR 2026
