AnthroTAP: Learning Point Tracking with Real-World Motion
Inès Hyeonsu Kim, Seokju Cho, Jahyeok Koo, Junghyun Park, Jiahui Huang, Honglak Lee, Joon-Young Lee, Seungryong Kim
摘要
Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly syntheticthe only source currently feasible to produce at scale. Collecting real-world annotations, however, is prohibitively expensive, as it requires tracking hundreds of points across frames. We introduce AnthroTAP, an automated pipeline that generates large-scale pseudo-labeled point tracking data from real human motion videos. Leveraging the structured complexity of human movementnon-rigid deformations, articulated motion, and frequent occlusionsAnthroTAP fits Skinned Multi-Person Linear (SMPL) models to detected humans, projects mesh vertices onto image planes, resolves occlusions via ray-casting, and filters unreliable tracks using optical flow consistency. A model trained on the AnthroTAP dataset achieves state-of-the-art performance on TAP-Vid, a challenging general-domain benchmark for tracking any point on diverse rigid and non-rigid objects (e.g., humans, animals, robots, and vehicles). Our approach outperforms recent self-training methods trained on vastly larger real datasets, while requiring only one day of training on 4 GPUs. AnthroTAP shows that structured human motion offers a scalable and effective source of real-world supervision for point tracking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MV-TAP: Tracking Any Point in Multi-View VideosJahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim, Junghyun Park 等CVPR 2026 · 被引用 4 次
- Real-World Point Tracking with Verifier-Guided Pseudo-LabelingGörkay Aydemir, Fatma Güney, Weidi XieCVPR 2026
- Generative Point Tracking and ForecastingXuanchen Lu, Ang Cao, Chao Feng, Andrew OwensCVPR 2026
它引用的顶会 Paper35
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa 等ICCV 2023 · 被引用 390 次
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan 等CVPR 2022 · 被引用 305 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
相关 Paper
- SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point TrackingWeiguang Zhao, Haoran Xu, Xingyu Miao, Qin Zhao 等SIGGRAPH 2026
- UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward PassMengfei Li, Peng Li, Zheng Zhang, Jiahao Lu 等CVPR 2026 · 被引用 7 次
- Generative Video MattingYongtao Ge, Kangyang Xie, Guangkai Xu, Li Ke 等SIGGRAPH 2025 · 被引用 1 次
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein 等ICCV 2023 · 被引用 255 次
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu 等CVPR 2020
