SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi, Mingxing Tan, Yin Cui, Quoc V. Le, Xiaodan Song
Abstract
Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scalepermuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by 3%+ AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.1% AP on COCO, attaining the new state-of-the-art performance for single model object detection without test-time augmentation. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/ tree/master/models/official/detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f2937fd-47bb-430d-ab7a-37ae50ee553dCited by top-tier papers26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui et al.NeurIPS 2020 · 755 citations
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai et al.NeurIPS 2021 · 228 citations
- PolyLoss: A Polynomial Expansion Perspective of Classification Loss FunctionsZhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk et al.ICLR 2022 · 189 citations
- GiraffeDet: A Heavy-Neck Paradigm for Object DetectionYiqi Jiang, Zhiyu Tan, Junyan Wang, Xiuyu Sun et al.ICLR 2022 · 147 citations
Builds on4
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 384 citations
- Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond ClassificationHang Xu, Lewei Yao, Zhenguo Li, Xiaodan Liang et al.ICCV 2019 · 197 citations
Related papers
- SP-NAS: Serial-to-Parallel Backbone Search for Object DetectionChenhan Jiang, Hang Xu, Wei Zhang, Xiaodan Liang et al.CVPR 2020
- RCNet: Reverse Feature Pyramid and Cross-scale Shift Network for Object DetectionZhuofan Zong, Qianggang Cao, Biao LengACM MM 2021 · 22 citations
- BFBox: Searching Face-Appropriate Backbone and Feature Pyramid Network for Face DetectorYang Liu, Xu TangCVPR 2020
- CBNet: A Novel Composite Backbone Network Architecture for Object DetectionYudong Liu, Yongtao Wang, Siwei Wang, Tingting Liang et al.AAAI 2020 · 266 citations
- Learning Rich Features at High-Speed for Single-Shot Object DetectionTiancai Wang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan et al.ICCV 2019 · 96 citations
