Track-On: Transformer-based Online Point Tracking with Memory
Görkay Aydemir, Xiongyi Cai, Weidi Xie, Fatma Güney
Abstract
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and occlusions. We target online tracking on a frame-by-frame basis, making it suitable for real-world, streaming scenarios. Specifically, we introduce Track-On, a simple transformer-based model designed for online long-term point tracking. Unlike prior methods that depend on full temporal modeling, our model processes video frames causally without access to future frames, leveraging two memory modules -- spatial memory and context memory -- to capture temporal information and maintain reliable point tracking over long time horizons. At inference time, it employs patch classification and refinement to identify correspondences and track points with high accuracy. Through extensive experiments, we demonstrate that Track-On sets a new state-of-the-art for online models and delivers superior or competitive results compared to offline approaches on seven datasets, including the TAP-Vid benchmark. Our method offers a robust and scalable solution for real-time tracking in diverse applications. Project page: https://kuis-ai.github.io/track_on
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AnthroTAP: Learning Point Tracking with Real-World MotionInès Hyeonsu Kim, Seokju Cho, Jahyeok Koo, Junghyun Park et al.CVPR 2026 · 5 citations
- Online Dense Point Tracking with Streaming MemoryQiaole Dong, Yanwei FuICCV 2025 · 1 citation
- Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkRulin Zhou, Wenlong He, An Wang, Jianhang Zhang et al.AAAI 2026
- Lattice Boltzmann Model for Learning Real-World Pixel DynamicityGuangze Zheng, Shijie Lin, Haobo Zuo, Si Si et al.NeurIPS 2025
- Generative Point Tracking and ForecastingXuanchen Lu, Ang Cao, Chao Feng, Andrew OwensCVPR 2026
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
Related papers
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
- Tapnext: Tracking Any Point (Tap) as Next Token PredictionArtem Zholus, Carl Doersch, Yi Yang, Skanda Koppula et al.ICCV 2025 · 7 citations
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long VideoJinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren et al.ICLR 2026 · 9 citations
- ReTracker: Exploring Image Matching for Robust Online Any Point TrackingDongli Tan, Xingyi He, Sida Peng, Yiqing Gong et al.ICCV 2025 · 2 citations
- Target-Aware Tracking with Long-Term Context AttentionKaijie He, Canlong Zhang, Sheng Xie, Zhixin Li et al.AAAI 2023 · 102 citations
