SpatialTracker: Tracking Any 2D Pixels in 3D Space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, Xiaowei Zhou
Abstract
Recovering dense and long-range pixel motion in videos is a challenging problem. Part of the difficulty arises from the 3D-to-2D projection process, leading to occlusions and discontinuities in the 2D motion domain. While 2D motion can be intricate, we posit that the underlying 3D motion can often be simple and low-dimensional. In this work, we propose to estimate point trajectories in 3D space to mitigate the issues caused by image projection. Our method, named SpatialTracker, lifts 2D pixels to 3D using monocular depth estimators, represents the 3D content of each frame efficiently using a triplane representation, and performs iterative updates using a transformer to estimate 3D trajectories. Tracking in 3D allows us to leverage asrigid-as-possible (ARAP) constraints while simultaneously learning a rigidity embedding that clusters pixels into different rigid parts. Extensive evaluation shows that our approach achieves state-of-the-art tracking performance both qualitatively and quantitatively, particularly in challenging scenarios such as out-of-plane rotation. And our project page is available at https://henry123-boy.github.io/SpaTracker/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b528661a-febb-4e91-9974-293bea4d8170Cited by top-tier papers9
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li et al.SIGGRAPH 2025 · 21 citations
- Fast Encoder-Based 3D from Casual Videos via Point Track ProcessingYoni Kasten, Wuyue Lu, Haggai MaronNeurIPS 2024 · 16 citations
- CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video GenerationQinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia et al.SIGGRAPH 2025 · 13 citations
- Dynamic View Synthesis as an Inverse ProblemHidir Yesiltepe, Pinar YanardagNeurIPS 2025 · 12 citations
- 3D-Fixup: Advancing Photo Editing with 3D PriorsYen-Chi Cheng, Krishna Kumar Singh, Jae Shin Yoon, Alexander G. Schwing et al.SIGGRAPH 2025 · 4 citations
Builds on16
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 314 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
Related papers
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev et al.ICCV 2025 · 6 citations
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsJiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li et al.NeurIPS 2025 · 7 citations
- TAPIP3D: Tracking Any Point in Persistent 3D GeometryBowei Zhang, Lei Ke, Adam W. Harley, Katerina FragkiadakiNeurIPS 2025 · 79 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
- KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in VideosDavid Novotný, Ignacio Rocco, Samarth Sinha, Alexandre Carlier et al.CVPR 2022 · 11 citations
