TAPIP3D: Tracking Any Point in Persistent 3D Geometry
Bowei Zhang, Lei Ke, Adam W. Harley, Katerina Fragkiadaki
摘要
We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camerastabilized spatio-temporal feature clouds, leveraging depth and camera motion information to lift 2D video features into a 3D world space where camera movement is effectively canceled out. Within this stabilized 3D representation, TAPIP3D iteratively refines multi-frame motion estimates, enabling robust point tracking over long time horizons. To handle the irregular structure of 3D point distributions, we propose a 3D Neighborhood-to-Neighborhood (N2N) attention mechanism-a 3Daware contextualization strategy that builds informative, spatially coherent feature neighborhoods to support precise trajectory estimation. Our 3D-centric formulation significantly improves performance over existing 3D point tracking methods and even surpasses state-of-the-art 2D pixel trackers in accuracy when reliable depth is available. The model supports inference in both camera-centric (unstabilized) and world-centric (stabilized) coordinates, with experiments showing that compensating for camera motion leads to substantial gains in tracking robustness. By replacing the conventional 2D square correlation windows used in prior 2D and 3D trackers with a spatially grounded 3D attention mechanism, TAPIP3D achieves strong and consistent results across multiple 3D point tracking benchmarks. Project Page: tapip3d.github.io
Camera pose sequence sampled at 4 time steps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta 等CVPR 2026 · 被引用 35 次
- DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed ImagesXiaoxue Chen, Ziyi Xiong, Yuantao Chen, Gen Li 等CVPR 2026 · 被引用 24 次
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang 等CVPR 2026 · 被引用 23 次
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment VideosSeungjae Lee, Yoonkyo Jung, Inkook Chun, Yao-Chih Lee 等CVPR 2026 · 被引用 17 次
- CoWTracker: Tracking by Warping instead of CorrelationZihang Lai, Eldar Insafutdinov, Edgar Sucar, Andrea VedaldiCVPR 2026 · 被引用 12 次
它引用的顶会 Paper29
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano 等CVPR 2022 · 被引用 984 次
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 被引用 314 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
相关 Paper
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev 等ICCV 2025 · 被引用 6 次
- SpatialTracker: Tracking Any 2D Pixels in 3D SpaceYuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue 等CVPR 2024 · 被引用 40 次
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long VideoJinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren 等ICLR 2026 · 被引用 9 次
- Context-PIPs: Persistent Independent Particles Demands Context FeaturesWeikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yitong Dong 等NeurIPS 2023 · 被引用 11 次
- MV-TAP: Tracking Any Point in Multi-View VideosJahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim, Junghyun Park 等CVPR 2026 · 被引用 4 次
