TrackIME: Enhanced Video Point Tracking via Instance Motion Estimation
Seong Hyeon Park, Huiwon Jang, Byungwoo Jeon, Sukmin Yun, Paul Hongsuck Seo, Jinwoo Shin
Abstract
Tracking points in video frames is essential for understanding video content. However, the task is fundamentally hindered by the computation demands for brute-force correspondence matching across the frames. As the current models down-sample the frame resolutions to mitigate this challenge, they fall short in accurately representing point trajectories due to information truncation. Instead, we address the challenge by pruning the search space for point tracking and let the model process only the important regions of the frames without down-sampling. Our first key idea is to identify the object instance and its trajectory over the frames, then prune the regions of the frame that do not contain the instance. Concretely, to estimate the instance’s trajectory, we track a group of points on the instance and aggregate their motion trajectories. Furthermore, to deal with the occlusions in complex scenes, we propose to compensate for the occluded points while tracking. To this end, we introduce a unified framework that jointly performs point tracking and segmentation, providing synergistic effects between the two tasks. For example, the segmentation results enable a tracking model to avoid the occluded points referring to the instance mask, and conversely, the improved tracking results can help to produce more accurate segmentation masks. Our framework can be easily incorporated with various tracking models, and we demonstrate its efficacy for enhanced point tracking throughout extensive experiments. For example, on the recent TAP-Vid benchmark, our framework consistently improves all baselines, e.g. , up to 13.5% improvement on the average Jaccard metric. The project url is https://trackime.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 942bfc9e-b91b-4c0b-a5da-984920c7e8fdBuilds on9
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
Related papers
- SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in VideosXiaowen Ying, Xin Li, Mooi Choo ChuahACM MM 2021 · 4 citations
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 18 citations
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi et al.AAAI 2020 · 80 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang et al.AAAI 2021 · 77 citations
