Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking
Jie Xiao, Yinchao Ma, Yuyang Tang, Dengqing Yang, Jianpeng Yang, Xu Zhou, Qiao Li, Wenfei Yang, Tianzhu Zhang
Abstract
3D single object tracking (SOT) in point clouds is essential for real-world 3D perception, yet it remains challenging due to data sparsity and large variations in scale and structure across diverse object categories. Most existing methods rely on a category-specific paradigm that trains separate models for each class, severely limiting scalability and generalization in real deployment. Extending these methods to a single model capable of tracking diverse object categories proves inadequate, as the significant variations across categories make it difficult to establish reliable geometric correspondences without category-specific priors. To overcome these limitations, we propose a Unified Structural KeyPoint Tracker (UniKPT), a novel structure-aware and generalizable framework for category-unified 3D point cloud tracking. UniKPT comprises three key modules: (1) an adaptive keypoint extractor that produces scale-aware and semantically meaningful keypoints; (2) a progressive correspondence aligner that establishes robust cross-frame geometric associations; and (3) a confidence-aware structural localization module that suppresses unreliable matches and leverages fine-grained structural relationships for precise 3D localization. Extensive experiments on the nuScenes and KITTI benchmarks show that UniKPT achieves new state-of-the-art performance in category-unified 3D SOT. On the challenging nuScenes dataset, our unified model further surpasses category-specific state-of-the-art trackers by 4.37 % in Success and 5.16 % in Precision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1250507b-3622-4a84-9bfc-ee3a08fed6feBuilds on26
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang et al.ICCV 2021 · 286 citations
- PTTR: Relational 3D Point Cloud Object Tracking with TransformerChangqing Zhou, Zhipeng Luo, Yueru Luo, Tianrui Liu et al.CVPR 2022 · 117 citations
Related papers
- PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D TrackingJuntao Liu, Zikun Zhou, Zhuotao Tian, Guangming Lu et al.ICLR 2026
- VoxelTrack: Exploring Multi-level Voxel Representation for 3D Point Cloud Object TrackingYuxuan Lu, Jiahao Nie, Zhiwei He, Hongjie Gu et al.ACM MM 2024 · 4 citations
- Towards Category Unification of 3D Single Object Tracking on Point CloudsJiahao Nie, Zhiwei He, Xudong Lv, Xueyi Zhou et al.ICLR 2024 · 20 citations
- TrackAny3D: Transferring Pretrained 3D Models for Category-Unified 3D Point Cloud TrackingMengmeng Wang, Haonan Wang, Yulong Li, Xiangjie Kong et al.ICCV 2025 · 2 citations
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildYifan Jiao, Yunhao Li, Junhua Ding, Qing Yang et al.ICCV 2025
