Fast Spatial Tracking with Visual Geometry Transformer
Chengjie Huang, GUILE WU, Dongfeng Bai, Bingbing Liu
摘要
Existing 3D point tracking methods mostly rely on heuristic designs or scene reconstruction, which incurs significant computational overhead and makes it difficult to meet the demands of real-time applications.To address this problem, in this work, we present VGGTracker, a novel spatial tracker that leverages a feed-forward visual geometry transformer to predict the trajectories of arbitrary query points from monocular videos in real time.Specifically, we employ a query initialization mechanism to maintain and update a global feature vector and a set of frame-level feature vectors for each query point.Then, we propose a new spatial tracking framework, which consists of a visual geometry transformer backbone, a global embedding branch, a frame-level embedding branch, and a tracking head.The key innovation lies in the dual-branch embedding design, where the global embedding branch integrates geometry-grounded features of the entire video into global query features to optimize track information across the entire sequence and the frame-level branch combines geometry-grounded features of each respective frame into frame-level query features to refine fine-grained track coordinate predictions.Furthermore, to facilitate collaboration between the global branch and the frame-level branch, we introduce an interaction module which enables unidirectional or bidirectional information exchange between the global query features and frame-level query features.Extensive experiments on various point tracking benchmark datasets show that our approach achieves significantly fast spatial tracking speed compared with state-of-the-art methods, while maintaining comparable tracking accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang 等ICLR 2026 · 被引用 318 次
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsRuicheng Wang, Sicheng Xu, Yue Dong, Yu Deng 等NeurIPS 2025 · 被引用 308 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein 等ICCV 2023 · 被引用 255 次
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch 等CVPR 2022 · 被引用 183 次
相关 Paper
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev 等ICCV 2025 · 被引用 6 次
- VideoTrack: Learning to Track Objects via Video TransformerFei Xie, Lei Chu, Jiahao Li, Yan Lu 等CVPR 2023
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等ICML 2026 · 被引用 12 次
- TrajVG: 3D Trajectory-Coupled Visual Geometry LearningXingyu Miao, Weiguang Zhao, Tao Lu, Linning Xu 等SIGGRAPH 2026
- KV-Tracker: Real-Time Pose Tracking with TransformersMarwan Taher, Ignacio Alzugaray, Kirill Mazur, Xin Kong 等CVPR 2026 · 被引用 5 次
