Bridging the Gap Between Detection and Tracking: A Unified Approach
Lianghua Huang, Xin Zhao, Kaiqi Huang
Abstract
Object detection models have been a source of inspiration for many tracking-by-detection algorithms over the past decade. Recent deep trackers borrow designs or modules from the latest object detection methods, such as bounding box regression, RPN and ROI pooling, and can deliver impressive performance. In this paper, instead of redesigning a new tracking-by-detection algorithm, we aim to explore a general framework for building trackers directly upon almost any advanced object detector. To achieve this, three key gaps must be bridged: (1) Object detectors are class-specific, while trackers are class-agnostic. (2) Object detectors do not differentiate intra-class instances, while this is a critical capability of a tracker. (3) Temporal cues are important for stable long-term tracking while they are not considered in still-image detectors. To address the above issues, we first present a simple target-guidance module for guiding the detector to locate target-relevant objects. Then a meta-learner is adopted for the detector to fast learn and adapt a target-distractor classifier online. We further introduce an anchored updating strategy to alleviate the problem of overfitting. The framework is instantiated on SSD [40] and FasterRCNN [15], the typical oneand two-stage detectors, respectively. Experiments on OTB, UAV123 and NfS have verified our framework and show that our trackers can benefit from deeper backbone networks, as opposed to many recent trackers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingLianghua Huang, Xin Zhao, Kaiqi HuangAAAI 2020 · 278 citations
- Global Tracking via Ensemble of Local TrackersZikun Zhou, Jianqiu Chen, Wenjie Pei, Kaige Mao et al.CVPR 2022 · 39 citations
- Effectiveness of Vision Transformer for Fast and Accurate Single-Stage Pedestrian DetectionJing Yuan, Panagiotis Barmpoutis, Tania StathakiNeurIPS 2022 · 11 citations
- Generalizable Pedestrian Detection: The Elephant in the RoomIrtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram et al.CVPR 2021
- Siam R-CNN: Visual Tracking by Re-DetectionPaul Voigtlaender, Jonathon Luiten, Philip H. S. Torr, Bastian LeibeCVPR 2020
Related papers
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training ModelBo Pang, Yizhuo Li, Yifan Zhang, Muchen Li et al.CVPR 2020
- Tracking by Instance Detection: A Meta-Learning ApproachGuangting Wang, Chong Luo, Xiaoyan Sun, Zhiwei Xiong et al.CVPR 2020
- Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template MatchingXuhua Huang, Jiarui Xu, Yu-Wing Tai, Chi-Keung TangCVPR 2020
- Generating Masks from Boxes by Mining Spatio-Temporal Consistencies in VideosBin Zhao, Goutam Bhat, Martin Danelljan, Luc Van Gool et al.ICCV 2021 · 19 citations
