Video Annotation for Visual Tracking via Selection and Refinement
Kenan Dai, Jie Zhao, Lijun Wang, Dong Wang, Jianhua Li, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang
Abstract
Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- Probabilistic Regression for Visual TrackingMartin Danelljan, Luc Van Gool, Radu TimofteCVPR 2020
- Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box EstimationBin Yan, Xinyu Zhang, Dong Wang, Huchuan Lu et al.CVPR 2021
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual TrackingDongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang et al.CVPR 2020
Related papers
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li et al.AAAI 2025 · 41 citations
- Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in VideosBin Zhao, Goutam Bhat, Martin Danelljan, Luc Van Gool et al.ICCV 2021 · 19 citations
- Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic SegmentationJungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee et al.ICCV 2019 · 45 citations
- Warp-Refine Propagation: Semi-Supervised Auto-labeling via Cycle-consistencyAditya Ganeshan, Alexis Vallet, Yasunori Kudo, Shin-ichi Maeda et al.ICCV 2021 · 14 citations
- Siamese Box Adaptive Network for Visual TrackingZedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang et al.CVPR 2020
