SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi, Mingxing Tan, Yin Cui, Quoc V. Le, Xiaodan Song
摘要
Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scalepermuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by 3%+ AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.1% AP on COCO, attaining the new state-of-the-art performance for single model object detection without test-time augmentation. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/ tree/master/models/official/detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai 等NeurIPS 2021 · 被引用 228 次
- PolyLoss: A Polynomial Expansion Perspective of Classification Loss FunctionsZhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk 等ICLR 2022 · 被引用 189 次
- GiraffeDet: A Heavy-Neck Paradigm for Object DetectionYiqi Jiang, Zhiyu Tan, Junyan Wang, Xiuyu Sun 等ICLR 2022 · 被引用 147 次
它引用的顶会 Paper4
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 被引用 384 次
- Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond ClassificationHang Xu, Lewei Yao, Zhenguo Li, Xiaodan Liang 等ICCV 2019 · 被引用 197 次
相关 Paper
- SP-NAS: Serial-to-Parallel Backbone Search for Object DetectionChenhan Jiang, Hang Xu, Wei Zhang, Xiaodan Liang 等CVPR 2020
- RCNet: Reverse Feature Pyramid and Cross-scale Shift Network for Object DetectionZhuofan Zong, Qianggang Cao, Biao LengACM MM 2021 · 被引用 22 次
- BFBox: Searching Face-Appropriate Backbone and Feature Pyramid Network for Face DetectorYang Liu, Xu TangCVPR 2020
- CBNet: A Novel Composite Backbone Network Architecture for Object DetectionYudong Liu, Yongtao Wang, Siwei Wang, Tingting Liang 等AAAI 2020 · 被引用 266 次
- Learning Rich Features at High-Speed for Single-Shot Object DetectionTiancai Wang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan 等ICCV 2019 · 被引用 96 次
