DriveTrack: A Benchmark for Long-Range Point Tracking in Real-World Videos
Arjun Balasingam, Joseph Chandler, Chenning Li, Zhoutong Zhang, Hari Balakrishnan
Abstract
This paper presents DriveTrack, a new benchmark and data generation framework for long-range keypoint tracking in real-world videos. DriveTrack is motivated by the observation that the accuracy of state-of-the-art trackers depends strongly on visual attributes around the selected keypoints, such as texture and lighting. The problem is that these artifacts are especially pronounced in real-world videos, but these trackers are unable to train on such scenes due to a dearth of annotations. DriveTrack bridges this gap by building a framework to automatically annotate point tracks on autonomous driving datasets. We release a dataset consisting of 1 billion point tracks across 24 hours of video, which is seven orders of magnitude greater than prior real-world benchmarks and on par with the scale of synthetic benchmarks. DriveTrack unlocks new use cases for point tracking in real-world videos. First, we show that finetuning keypoint trackers on DriveTrack improves accuracy on real-world scenes by up to 7%. Second, we analyze the sensitivity of trackers to visual artifacts in real scenes and motivate the idea of running assistive keypoint selectors alongside trackers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- TAPIP3D: Tracking Any Point in Persistent 3D GeometryBowei Zhang, Lei Ke, Adam W. Harley, Katerina FragkiadakiNeurIPS 2025 · 79 citations
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen et al.CVPR 2026 · 64 citations
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- AllTracker: Efficient Dense Point Tracking at High ResolutionAdam W. Harley, Yang You, Xinglong Sun, Yang Zheng et al.ICCV 2025 · 8 citations
- Tapnext: Tracking Any Point (Tap) as Next Token PredictionArtem Zholus, Carl Doersch, Yi Yang, Skanda Koppula et al.ICCV 2025 · 7 citations
Builds on7
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard et al.CVPR 2020
Related papers
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingYang Zhou, Hao Shao, Letian Wang, Zhuofan Zong et al.ICLR 2026 · 20 citations
- PlanarTrack: A Large-scale Challenging Benchmark for Planar Object TrackingXinran Liu, Xiaoqiong Liu, Ziruo Yi, Xin Zhou et al.ICCV 2023 · 2 citations
- SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic AnnotationsYunnan Wang, Kecheng Zheng, Jianyuan Wang, Minghao Chen et al.CVPR 2026 · 1 citation
- Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingJiawei Ge, Xinyu Zhang, Jiuxin Cao, Xuelin Zhu et al.ACM MM 2025 · 3 citations
- Real-World Point Tracking with Verifier-Guided Pseudo-LabelingGörkay Aydemir, Fatma Güney, Weidi XieCVPR 2026
