MobileOne: An Improved One millisecond Mobile Backbone
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, Anurag Ranjan
摘要
Efficient neural network backbones for mobile devices are often optimized for metrics such as FLOPs or parameter count. However, these metrics may not correlate well with latency of the network when deployed on a mobile device. Therefore, we perform extensive analysis of different metrics by deploying several mobile-friendly networks on a mobile device. We identify and analyze architectural and optimization bottlenecks in recent efficient neural networks and provide ways to mitigate these bottlenecks. To this end, we design an efficient backbone MobileOne, with variants achieving an inference time under 1 ms on an iPhone12 with 75.9% top-1 accuracy on ImageNet. We show that Mo-bileOne achieves state-of-the-art performance within the efficient architectures while being many times faster on mobile. Our best model obtains similar performance on Ima-geNet as MobileFormer while being 38× faster. Our model obtains 2.3% better top-1 accuracy on ImageNet than Ef-ficientNet at similar latency. Furthermore, we show that our model generalizes to multiple tasks -image classification, object detection, and semantic segmentation with significant improvements in latency and accuracy as compared to existing efficient architectures when deployed on a mobile device. Code and models are available at https: //github.com/apple/ml-mobileone
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han 等CVPR 2024 · 被引用 500 次
- SHViT: Single-Head Vision Transformer with Memory Efficient Macro DesignSeokju Yun, Youngmin RoCVPR 2024 · 被引用 117 次
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim 等NeurIPS 2023 · 被引用 109 次
- SeaFormer: Squeeze-enhanced Axial Transformer for Mobile Semantic SegmentationQiang Wan, Zilong Huang, Jiachen Lu, Gang Yu 等ICLR 2023 · 被引用 82 次
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao 等ICLR 2024 · 被引用 30 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
相关 Paper
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
- Iformer: Integrating ConvNet and Transformer for Mobile ApplicationChuanyang ZhengICLR 2025
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel 等ICCV 2023 · 被引用 341 次
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu 等CVPR 2022 · 被引用 600 次
- MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile DevicesBo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin 等CVPR 2020
