Scale-Equalizing Pyramid Convolution for Object Detection
Xinjiang Wang, Shilong Zhang, Zhuoran Yu, Litong Feng, Wayne Zhang
摘要
Feature pyramid has been an efficient method to extract features at different scales. Development over this method mainly focuses on aggregating contextual information at different levels while seldom touching the inter-level correlation in the feature pyramid. Early computer vision methods extracted scale-invariant features by locating the feature extrema in both spatial and scale dimension. Inspired by this, a convolution across the pyramid level is proposed in this study, which is termed pyramid convolution and is a modified 3-D convolution. Stacked pyramid convolutions directly extract 3-D (scale and spatial) features and outperforms other meticulously designed feature fusion modules. Based on the viewpoint of 3-D convolution, an integrated batch normalization that collects statistics from the whole feature pyramid is naturally inserted after the pyramid convolution. Furthermore, we also show that the naive pyramid convolution, together with the design of RetinaNet head, actually best applies for extracting features from a Gaussian pyramid, whose properties can hardly be satisfied by a feature pyramid. In order to alleviate this discrepancy, we build a scale-equalizing pyramid convolution (SEPC) that aligns the shared pyramid convolution kernel only at high-level feature maps. Being computationally efficient and compatible with the head design of most single-stage object detectors, the SEPC module brings significant performance improvement (> 4AP increase on MS-COCO2017 dataset) in state-of-the-art one-stage object detectors, and a light version of SEPC also has ∼ 3.5AP gain with only around 7% inference time increase. The pyramid convolution also functions well as a stand-alone module in two-stage object detectors and is able to improve the performance by ∼ 2AP. The source code can be found at https://github.com/jshilong/SEPC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object DetectionChenhongyi Yang, Zehao Huang, Naiyan WangCVPR 2022 · 被引用 472 次
- Dynamic DETR: End-to-End Object Detection with Dynamic AttentionXiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang 等ICCV 2021 · 被引用 429 次
- RCNet: Reverse Feature Pyramid and Cross-scale Shift Network for Object DetectionZhuofan Zong, Qianggang Cao, Biao LengACM MM 2021 · 被引用 22 次
- DM-EFS: Dynamically Multiplexed Expanded Features Set form for Robust and Efficient Small Object DetectionAashish SharmaICCV 2025 · 被引用 2 次
- Scaled-YOLOv4: Scaling Cross Stage Partial NetworkChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2021
它引用的顶会 Paper5
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang 等ICCV 2019 · 被引用 1,056 次
- Scale-Aware Trident Networks for Object DetectionYanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 1,031 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
相关 Paper
- FAS-Net: Construct Effective Features Adaptively for Multi-Scale Object DetectionJiangqiao Yan, Yue Zhang, Zhonghan Chang, Tengfei Zhang 等AAAI 2020 · 被引用 2 次
- AugFPN: Improving Multi-Scale Feature Learning for Object DetectionChaoxu Guo, Bin Fan, Qian Zhang, Shiming Xiang 等CVPR 2020
- You Only Look One-Level FeatureQiang Chen, Yingming Wang, Tong Yang, Xiangyu Zhang 等CVPR 2021
- GraphFPN: Graph Feature Pyramid Network for Object DetectionGangming Zhao, Weifeng Ge, Yizhou YuICCV 2021 · 被引用 122 次
- Construct Effective Geometry Aware Feature Pyramid Network for Multi-Scale Object DetectionJinpeng Dong, Yuhao Huang, Songyi Zhang, Shitao Chen 等AAAI 2022 · 被引用 9 次
