Learning Rich Features at High-Speed for Single-Shot Object Detection
Tiancai Wang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao
Abstract
Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageNet for the classification task and use a top-down feature pyramid representation for handling scale variations. Contrary to common pre-training strategy, recent works have demonstrated the benefits of training from scratch to reduce the task gap between classification and localization, especially at high overlap thresholds. However, detection models trained from scratch require significantly longer training time compared to their typical finetuning based counterparts. We introduce a single-stage detection framework that combines the advantages of both fine-tuning pretrained models and training from scratch. Our framework constitutes a standard network that uses a pre-trained backbone and a parallel light-weight auxiliary network trained from scratch. Further, we argue that the commonly used top-down pyramid representation only focuses on passing high-level semantics from the top layers to bottom layers. We introduce a bi-directional network that efficiently circulates both low-/mid-level and high-level semantic information in the detection framework. Experiments are performed on MS COCO and UAVDT datasets. Compared to the baseline, our detector achieives an absolute gain of 7.4% and 4.2% in average precision (AP) on MS COCO and UAVDT datasets, respectively using VGG backbone. For a 300×300 input on the MS COCO test set, our detector with ResNet backbone surpasses existing single-stage detection methods for single-scale inference achieving 34.3 AP, while operating at an inference time of 19 milliseconds on a single Titan X GPU. Code is avail- able at https://github.com/vaesl/LRF-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67fd736e-4311-421b-ac03-39ef4d0eabedCited by top-tier papers5
- Hierarchical Shot DetectorJiale Cao, Yanwei Pang, Jungong Han, Xuelong LiICCV 2019 · 70 citations
- Co-mining: Self-Supervised Learning for Sparsely Annotated Object DetectionTiancai Wang, Tong Yang, Jiale Cao, Xiangyu ZhangAAAI 2021 · 57 citations
- Learning Human-Object Interaction Detection Using Interaction PointsTiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan et al.CVPR 2020
- Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample SelectionShifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei et al.CVPR 2020
- D2Det: Towards High Quality Object Detection and Instance SegmentationJiale Cao, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan et al.CVPR 2020
Builds on2
Related papers
- RCNet: Reverse Feature Pyramid and Cross-scale Shift Network for Object DetectionZhuofan Zong, Qianggang Cao, Biao LengACM MM 2021 · 22 citations
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang et al.CVPR 2022 · 14 citations
- Enriched Feature Guided Refinement Network for Object DetectionJing Nie, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan et al.ICCV 2019 · 80 citations
- SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationXianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi et al.CVPR 2020
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object DetectionLewei Yao, Hang Xu, Wei Zhang, Xiaodan Liang et al.AAAI 2020 · 83 citations
