TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
Jinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren, Zhaoyang Zeng, Lei Zhang
Abstract
In this paper, built upon TAPTRv2, we present TAPTRv3. TAPTRv2 is a simple yet effective DETR-like point tracking framework that works fine in regular videos but tends to fail in long videos. TAPTRv3 improves TAPTRv2 by addressing its shortcomings in querying high-quality features from long videos, where the target tracking points normally undergo increasing variation over time. In TAPTRv3, we propose to utilize both spatial and temporal context to bring better feature querying along the spatial and temporal dimensions for more robust tracking in long videos. For better spatial feature querying, we identify that off-the-shelf attention mechanisms struggle with point-level tasks and present Context-aware Cross-Attention (CCA). CCA introduces spatial context into the attention mechanism to enhance the quality of attention scores when querying image features. For better temporal feature querying, we introduce Visibility-aware Long-Temporal Attention (VLTA), which conducts temporal attention over past frames while considering their corresponding visibilities. This effectively addresses the feature drifting problem in TAPTRv2 caused by its RNN-like long-term modeling. TAPTRv3 surpasses TAPTRv2 by a large margin on most of the challenging datasets and obtains state-of-the-art performance. Even when compared with methods trained on large-scale extra internal data, TAPTRv3 still demonstrates superiority. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Tapnext: Tracking Any Point (Tap) as Next Token PredictionArtem Zholus, Carl Doersch, Yi Yang, Skanda Koppula et al.ICCV 2025 · 7 citations
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsJiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li et al.NeurIPS 2025 · 7 citations
- AnthroTAP: Learning Point Tracking with Real-World MotionInès Hyeonsu Kim, Seokju Cho, Jahyeok Koo, Junghyun Park et al.CVPR 2026 · 5 citations
- Fast Spatial Tracking with Visual Geometry TransformerChengjie Huang, GUILE WU, Dongfeng Bai, Bingbing LiuCVPR 2026
- E-MaT: Event-oriented Mamba for Egocentric Point TrackingHan Han, Wei Zhai, Baocai Yin, Yang Cao et al.AAAI 2026
Builds on19
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
Related papers
- TAPTRv2: Attention-based Position Update Improves Tracking Any PointHongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng et al.NeurIPS 2024 · 22 citations
- Learning Spatial-Semantic Features for Robust Video Object SegmentationXin Li, Deshui Miao, Zhenyu He, Yaowei Wang et al.ICLR 2025
- TAPIP3D: Tracking Any Point in Persistent 3D GeometryBowei Zhang, Lei Ke, Adam W. Harley, Katerina FragkiadakiNeurIPS 2025 · 79 citations
- Context-PIPs: Persistent Independent Particles Demands Context FeaturesWeikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yitong Dong et al.NeurIPS 2023 · 11 citations
- Track-On: Transformer-based Online Point Tracking with MemoryGörkay Aydemir, Xiongyi Cai, Weidi Xie, Fatma GüneyICLR 2025
