Video Annotation for Visual Tracking via Selection and Refinement
Kenan Dai, Jie Zhao, Lijun Wang, Dong Wang, Jianhua Li, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang
摘要
Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan 等AAAI 2020 · 被引用 944 次
- Probabilistic Regression for Visual TrackingMartin Danelljan, Luc Van Gool, Radu TimofteCVPR 2020
- Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box EstimationBin Yan, Xinyu Zhang, Dong Wang, Huchuan Lu 等CVPR 2021
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual TrackingDongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang 等CVPR 2020
相关 Paper
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2025 · 被引用 41 次
- Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in VideosBin Zhao, Goutam Bhat, Martin Danelljan, Luc Van Gool 等ICCV 2021 · 被引用 19 次
- Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic SegmentationJungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee 等ICCV 2019 · 被引用 45 次
- Warp-Refine Propagation: Semi-Supervised Auto-labeling via Cycle-consistencyAditya Ganeshan, Alexis Vallet, Yasunori Kudo, Shin-ichi Maeda 等ICCV 2021 · 被引用 14 次
- Siamese Box Adaptive Network for Visual TrackingZedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang 等CVPR 2020
