TrackIME: Enhanced Video Point Tracking via Instance Motion Estimation
Seong Hyeon Park, Huiwon Jang, Byungwoo Jeon, Sukmin Yun, Paul Hongsuck Seo, Jinwoo Shin
摘要
Tracking points in video frames is essential for understanding video content. However, the task is fundamentally hindered by the computation demands for brute-force correspondence matching across the frames. As the current models down-sample the frame resolutions to mitigate this challenge, they fall short in accurately representing point trajectories due to information truncation. Instead, we address the challenge by pruning the search space for point tracking and let the model process only the important regions of the frames without down-sampling. Our first key idea is to identify the object instance and its trajectory over the frames, then prune the regions of the frame that do not contain the instance. Concretely, to estimate the instance’s trajectory, we track a group of points on the instance and aggregate their motion trajectories. Furthermore, to deal with the occlusions in complex scenes, we propose to compensate for the occluded points while tracking. To this end, we introduce a unified framework that jointly performs point tracking and segmentation, providing synergistic effects between the two tasks. For example, the segmentation results enable a tracking model to avoid the occluded points referring to the instance mask, and conversely, the improved tracking results can help to produce more accurate segmentation masks. Our framework can be easily incorporated with various tracking models, and we demonstrate its efficacy for enhanced point tracking throughout extensive experiments. For example, on the recent TAP-Vid benchmark, our framework consistently improves all baselines, e.g. , up to 13.5% improvement on the average Jaccard metric. The project url is https://trackime.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 等NeurIPS 2023 · 被引用 709 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein 等ICCV 2023 · 被引用 255 次
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li 等ICCV 2023 · 被引用 238 次
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch 等CVPR 2022 · 被引用 183 次
相关 Paper
- SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in VideosXiaowen Ying, Xin Li, Mooi Choo ChuahACM MM 2021 · 被引用 4 次
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 被引用 18 次
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi 等AAAI 2020 · 被引用 80 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang 等AAAI 2021 · 被引用 77 次
