MobileOne: An Improved One millisecond Mobile Backbone
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, Anurag Ranjan
Abstract
Efficient neural network backbones for mobile devices are often optimized for metrics such as FLOPs or parameter count. However, these metrics may not correlate well with latency of the network when deployed on a mobile device. Therefore, we perform extensive analysis of different metrics by deploying several mobile-friendly networks on a mobile device. We identify and analyze architectural and optimization bottlenecks in recent efficient neural networks and provide ways to mitigate these bottlenecks. To this end, we design an efficient backbone MobileOne, with variants achieving an inference time under 1 ms on an iPhone12 with 75.9% top-1 accuracy on ImageNet. We show that Mo-bileOne achieves state-of-the-art performance within the efficient architectures while being many times faster on mobile. Our best model obtains similar performance on Ima-geNet as MobileFormer while being 38× faster. Our model obtains 2.3% better top-1 accuracy on ImageNet than Ef-ficientNet at similar latency. Furthermore, we show that our model generalizes to multiple tasks -image classification, object detection, and semantic segmentation with significant improvements in latency and accuracy as compared to existing efficient architectures when deployed on a mobile device. Code and models are available at https: //github.com/apple/ml-mobileone
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe8c3db5-6fd8-4c77-8883-b1885388ea2aCited by top-tier papers36
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han et al.CVPR 2024 · 500 citations
- SHViT: Single-Head Vision Transformer with Memory Efficient Macro DesignSeokju Yun, Youngmin RoCVPR 2024 · 117 citations
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim et al.NeurIPS 2023 · 109 citations
- SeaFormer: Squeeze-enhanced Axial Transformer for Mobile Semantic SegmentationQiang Wan, Zilong Huang, Jiachen Lu, Gang Yu et al.ICLR 2023 · 82 citations
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICLR 2024 · 30 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
Related papers
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
- Iformer: Integrating ConvNet and Transformer for Mobile ApplicationChuanyang ZhengICLR 2025
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.ICCV 2023 · 341 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
- MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile DevicesBo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin et al.CVPR 2020
