Online Dense Point Tracking with Streaming Memory
Qiaole Dong, Yanwei Fu
Abstract
Dense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence of occlusions. Traditional methods use optical flow models to directly estimate long-range motion, but they often suffer from appearance drifting without considering temporal consistency. Recent point tracking algorithms usually depend on sliding windows for indirect information propagation from the first frame to the current one, which is slow and less effective for long-range tracking. To account for temporal consistency and enable efficient information propagation, we present a lightweight and fast model with Streaming memory for dense POint Tracking and online video processing. The SPOT framework features three core components: a customized memory reading module for feature enhancement, a sensory memory for short-term motion dynamics modeling, and a visibility-guided splatting module for accurate information propagation. This combination enables SPOT to perform dense point tracking with state-of-the-art accuracy on the CVO benchmark, as well as comparable or superior performance to offline models on sparse tracking benchmarks such as TAP-Vid and RoboTAP. Notably, SPOT with smaller parameter numbers operates at least faster than previous state-of-the-art models while maintaining the best performance on CVO. We will release the models and codes at: https://dqiaole.github.io/SPOT/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6229084-0dc3-44dc-9913-5bd66f36372bCited by top-tier papers1
Ask how each one uses itBuilds on25
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
Related papers
- Dense Optical Tracking: Connecting the DotsGuillaume Le Moing, Jean Ponce, Cordelia SchmidCVPR 2024 · 24 citations
- CoWTracker: Tracking by Warping instead of CorrelationZihang Lai, Eldar Insafutdinov, Edgar Sucar, Andrea VedaldiCVPR 2026 · 12 citations
- Track-On: Transformer-based Online Point Tracking with MemoryGörkay Aydemir, Xiongyi Cai, Weidi Xie, Fatma GüneyICLR 2025
- TrackIME: Enhanced Video Point Tracking via Instance Motion EstimationSeong Hyeon Park, Huiwon Jang, Byungwoo Jeon, Sukmin Yun et al.NeurIPS 2024 · 1 citation
- AllTracker: Efficient Dense Point Tracking at High ResolutionAdam W. Harley, Yang You, Xinglong Sun, Yang Zheng et al.ICCV 2025 · 8 citations
