ThunderNet: Towards Real-Time Generic Object Detection on Mobile Devices
Zheng Qin, Zeming Li, Zhaoning Zhang, Yiping Bao, Gang Yu, Yuxing Peng, Jian Sun
Abstract
Real-time generic object detection on mobile platforms is a crucial but challenging computer vision task. Prior lightweight CNN-based detectors are inclined to use onestage pipeline. In this paper, we investigate the effectiveness of two-stage detectors in real-time generic detection and propose a lightweight two-stage detector named Thun-derNet. In the backbone part, we analyze the drawbacks in previous lightweight backbones and present a lightweight backbone designed for object detection. In the detection part, we exploit an extremely efficient RPN and detection head design. To generate more discriminative feature representation, we design two efficient architecture blocks, Context Enhancement Module and Spatial Attention Module. At last, we investigate the balance between the input resolution, the backbone, and the detection head. Benefit from the highly efficient backbone and detection part design, ThunderNet surpasses previous lightweight one-stage detectors with only 40% of the computational cost on PAS-CAL VOC and COCO benchmarks. Without bells and whistles, ThunderNet runs at 24.1 fps on an ARM-based device with 19.2 AP on COCO. To the best of our knowledge, this is the first real-time detector reported on ARM platforms. Our code and models are available at https: //github.com/qinzheng93/ThunderNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aff7511-5c98-4153-861d-84c9c0e0b42eCited by top-tier papers15
- Rethinking Transformer-based Set Prediction for Object DetectionZhiqing Sun, Shengcao Cao, Yiming Yang, Kris KitaniICCV 2021 · 381 citations
- Residual Attention: A Simple but Effective Method for Multi-Label RecognitionKe Zhu, Jianxin WuICCV 2021 · 190 citations
- Task-Oriented Feature DistillationLinfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma et al.NeurIPS 2020 · 74 citations
- Q-VLM: Post-training Quantization for Large Vision-Language ModelsChangyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang et al.NeurIPS 2024 · 57 citations
- Multi-Source Domain Adaptation for Object DetectionXingxu Yao, Sicheng Zhao, Pengfei Xu, Jufeng YangICCV 2021 · 54 citations
Related papers
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- YOLO-ULM: Ultra-Lightweight Models for Real-Time Object DetectionShasha Han, Chong Li, Xinning Wang, Xuebo LiCVPR 2026
- Oriented R-CNN for Object DetectionXingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao et al.ICCV 2021 · 1,070 citations
- Computation Reallocation for Object DetectionFeng Liang, Chen Lin, Ronghao Guo, Ming Sun et al.ICLR 2020 · 36 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
