TrackAny3D: Transferring Pretrained 3D Models for Category-Unified 3D Point Cloud Tracking
Mengmeng Wang, Haonan Wang, Yulong Li, Xiangjie Kong, Jiaxin Du, Guojiang Shen, Feng Xia
摘要
3D LiDAR-based single object tracking (SOT) relies on sparse and irregular point clouds, posing challenges from geometric variations in scale, motion patterns, and structural complexity across object categories. Current category-specific approaches achieve good accuracy but are impractical for real-world use, requiring separate models for each category and showing limited generalization. To tackle these issues, we propose TrackAny3D, the first framework to transfer large-scale pretrained 3D models for category-agnostic 3D SOT. We first integrate parameterefficient adapters to bridge the gap between pretraining and tracking tasks while preserving geometric priors. Then, we introduce a Mixture-of-Geometry-Experts (MoGE) architecture that adaptively activates specialized subnetworks based on distinct geometric characteristics. Additionally, we design a temporal context optimization strategy that incorporates learnable temporal tokens and a dynamic mask weighting module to propagate historical information and mitigate temporal drift. Experiments on three commonlyused benchmarks show that TrackAny3D establishes new state-of-the-art performance on category-agnostic 3D SOT, demonstrating strong generalization and competitiveness. We hope this work will enlighten the community on the importance of unified models and further expand the use of large-scale pretrained models in this field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D TrackingJuntao Liu, Zikun Zhou, Zhuotao Tian, Guangming Lu 等ICLR 2026
- Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object TrackingJie Xiao, Yinchao Ma, Yuyang Tang, Dengqing Yang 等CVPR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
相关 Paper
- Towards Category Unification of 3D Single Object Tracking on Point CloudsJiahao Nie, Zhiwei He, Xudong Lv, Xueyi Zhou 等ICLR 2024 · 被引用 20 次
- M3SOT: Multi-Frame, Multi-Field, Multi-Space 3D Single Object TrackingJiaming Liu, Yue Wu, Maoguo Gong, Qiguang Miao 等AAAI 2024 · 被引用 17 次
- VoxelTrack: Exploring Multi-level Voxel Representation for 3D Point Cloud Object TrackingYuxuan Lu, Jiahao Nie, Zhiwei He, Hongjie Gu 等ACM MM 2024 · 被引用 4 次
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildYifan Jiao, Yunhao Li, Junhua Ding, Qing Yang 等ICCV 2025
- Modeling Continuous Motion for 3D Point Cloud Object TrackingZhipeng Luo, Gongjie Zhang, Changqing Zhou, Zhonghua Wu 等AAAI 2024 · 被引用 10 次
