Rethinking Mobile Block for Efficient Attention-based Models
Jiangning Zhang, Xiangtai Li, Jian Li, Liang Liu, Zhucun Xue, Boshen Zhang, Zhengkai Jiang, Tianxin Huang, Yabiao Wang, Chengjie Wang
Abstract
This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This work rethinks lightweight infrastructure from efficient IRB and effective components of Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMB) for lightweight model design. Following simple but effective design criterion, we deduce a modern Inverted Residual Mobile Block (iRMB) and build a ResNetlike Efficient MOdel (EMO) with only iRMB for down-stream tasks. Extensive experiments on ImageNet-1K, COCO2017, and ADE20K benchmarks demonstrate the superiority of our EMO over state-of-the-art methods, e.g., EMO-1M/2M/5M achieve 71.5, 75.1, and 78.4 Top-1 that surpass equal-order CNN-/Attention-based models, while trading-off the parameter, efficiency, and accuracy well: running 2.8-4.0× ↑ faster than EdgeNeXt on iPhone14.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69574e87-3270-4f9d-bb73-d9adbd5b653eCited by top-tier papers18
- A Diffusion-Based Framework for Multi-Class Anomaly DetectionHaoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen et al.AAAI 2024 · 231 citations
- Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelZeyu Wang, Chen Li, Huiying Xu, Xinzhong Zhu et al.AAAI 2025 · 136 citations
- SHViT: Single-Head Vision Transformer with Memory Efficient Macro DesignSeokju Yun, Youngmin RoCVPR 2024 · 117 citations
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationFenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma et al.ACM MM 2025 · 29 citations
- DGMamba: Domain Generalization via Generalized State Space ModelShaocong Long, Qianyu Zhou, Xiangtai Li, Xuequan Lu et al.ACM MM 2024 · 15 citations
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision ModelsChenglin Yang, Siyuan Qiao, Qihang Yu, Xiaoding Yuan et al.ICLR 2023 · 22 citations
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han et al.CVPR 2024 · 500 citations
- Edge-oriented Convolution Block for Real-time Super Resolution on Mobile DevicesXindong Zhang, Hui Zeng, Lei ZhangACM MM 2021 · 229 citations
- MobileOne: An Improved One millisecond Mobile BackbonePavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.CVPR 2023
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICLR 2024 · 30 citations
