Tracking Everything Everywhere across Multiple Cameras
Li-Heng Wang, YuJu Cheng, Tyng-Luh Liu
Abstract
Pixel tracking in single-view video sequences has recently emerged as a significant area of research. While previous work has primarily concentrated on tracking within a given video, we propose to expand pixel correspondence estimation into multi-view scenarios. The central concept involves utilizing a canonical space that preserves a universal 3D representation across different views and timesteps. This model allows for precise tracking of points even through prolonged occlusions and significant deformations in appearance between views. Moreover, we show that our model, through the use of an efficient training strategy incorporating distillation loss, is capable of performing incremental pixel tracking, a process often seen as complex in test-time optimization techniques. Comprehensive experiments validate the method's ability to accurately establish point correspondences across cameras. Furthermore, our method achieves promising results of multi-view pixel tracking without requiring the entire video sequences to be provided at once.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Neural 3D Video Synthesis from Multi-view VideoTianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green et al.CVPR 2022 · 324 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
- Immersive light field video with a layered mesh representationMichael Broxton, John Flynn, Ryan S. Overbeck, Daniel Erickson et al.SIGGRAPH 2020 · 271 citations
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li et al.ICCV 2023 · 112 citations
Related papers
- Seeing Behind Objects for 3D Multi-Object Tracking in RGB-D SequencesNorman Müller, Yu-Shiang Wong, Niloy J. Mitra, Angela Dai et al.CVPR 2021
- Delta: Dense Efficient Long-Range 3D tracking for any videoTuan Duc Ngo, Peiye Zhuang, Evangelos Kalogerakis, Chuang Gan et al.ICLR 2025
- SpatialTracker: Tracking Any 2D Pixels in 3D SpaceYuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue et al.CVPR 2024 · 40 citations
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstructionDavid Novotný, Roman Shapovalov, Andrea VedaldiNeurIPS 2020 · 9 citations
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev et al.ICCV 2025 · 6 citations
