SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera Motion
Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Yuri Makarov, Bingyi Kang, Xing Zhu, Hujun Bao, Yujun Shen, Xiaowei Zhou
摘要
We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50x faster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco 等CVPR 2026 · 被引用 52 次
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationSixu Lin, Junliang Chen, Huaiyuan Xu, Zhuohao Li 等ICML 2026 · 被引用 3 次
- FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction VideosAlexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario 等CVPR 2026
- MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid CamerasYiqian Chang, Qinghong Ye, Haoran Xu, Jianing Li 等CVPR 2026
- MLLM-4D: Towards Visual-based Spatial-Temporal IntelligenceXingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun 等ICML 2026
它引用的顶会 Paper45
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
相关 Paper
- SpatialTracker: Tracking Any 2D Pixels in 3D SpaceYuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue 等CVPR 2024 · 被引用 40 次
- ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the WorldHaiwen Feng, Junyi Zhang, Qianqian Wang, Yufei Ye 等ICCV 2025 · 被引用 14 次
- Fast Spatial Tracking with Visual Geometry TransformerChengjie Huang, GUILE WU, Dongfeng Bai, Bingbing LiuCVPR 2026
- Track3R: Joint Point Map and Trajectory Prior for Spatiotemporal 3D UnderstandingSeong Hyeon Park, Jinwoo ShinNeurIPS 2025
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等ICML 2026 · 被引用 12 次
