Searching for MobileNetV3
Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, Yukun Zhu
Abstract
We present the next generation of MobileNets based on a combination of complementary search techniques as well as a novel architecture design. MobileNetV3 is tuned to mobile phone CPUs through a combination of hardware-aware network architecture search (NAS) complemented by the NetAdapt algorithm and then subsequently improved through novel architecture advances. This paper starts the exploration of how automated search algorithms and network design can work together to harness complementary approaches improving the overall state of the art. Through this process we create two new MobileNet models for release: MobileNetV3-Large and MobileNetV3-Small which are targeted for high and low resource use cases. These models are then adapted and applied to the tasks of object detection and semantic segmentation. For the task of semantic segmentation (or any dense pixel prediction), we propose a new efficient segmentation decoder Lite Reduced Atrous Spatial Pyramid Pooling (LR-ASPP). We achieve new state of the art results for mobile classification, detection and segmentation. MobileNetV3-Large is 3.2% more accurate on ImageNet classification while reducing latency by 20% compared to MobileNetV2. MobileNetV3-Small is 6.6% more accurate compared to a MobileNetV2 model with comparable latency. MobileNetV3-Large detection is over 25% faster at roughly the same accuracy as MobileNetV2 on COCO detection. MobileNetV3-Large LR-ASPP is 34% faster than MobileNetV2 R-ASPP at similar accuracy for Cityscapes segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2566c049-9733-4b1d-8fbd-2e2a034746b8Cited by top-tier papers493
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural NetworksLingxiao Yang, Ru-Yuan Zhang, Lida Li, Xiaohua XieICML 2021 · 1,593 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 1,298 citations
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock et al.ICCV 2021 · 1,009 citations
Related papers
- MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile DevicesBo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin et al.CVPR 2020
- Fast Neural Network Adaptation via Parameter Remapping and Architecture SearchJiemin Fang, Yuzhu Sun, Kangjian Peng, Qian Zhang et al.ICLR 2020 · 36 citations
- MobileDets: Searching for Object Detection Architectures for Mobile AcceleratorsYunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin et al.CVPR 2021
- Computation Reallocation for Object DetectionFeng Liang, Chen Lin, Ronghao Guo, Ming Sun et al.ICLR 2020 · 36 citations
- SP-NAS: Serial-to-Parallel Backbone Search for Object DetectionChenhan Jiang, Hang Xu, Wei Zhang, Xiaodan Liang et al.CVPR 2020
