Iformer: Integrating ConvNet and Transformer for Mobile Application
Chuanyang Zheng
Abstract
We present a new family of mobile hybrid vision networks, called iFormer, with a focus on optimizing latency and accuracy on mobile applications. iFormer effectively integrates the fast local representation capacity of convolution with the efficient global modeling ability of self-attention. The local interactions are derived from transforming a standard convolutional network, i.e., ConvNeXt, to design a more lightweight mobile network. Our newly introduced mobile modulation attention removes memory-intensive operations in MHA and employs an efficient modulation mechanism to boost dynamic global representational capacity. We conduct comprehensive experiments demonstrating that iFormer outperforms existing lightweight networks across various tasks. Notably, iFormer achieves an impressive Top-1 accuracy of 80.4% on ImageNet-1k with a latency of only 1.10 ms on an iPhone 13, surpassing the recently proposed MobileNetV4 under similar latency constraints. Additionally, our method shows significant improvements in downstream tasks, including COCO object detection, instance segmentation, and ADE20k semantic segmentation, while still maintaining low latency on mobile devices for high-resolution inputs in these scenarios. Code and models are available at: https://github.com/ChuanyangZheng/iFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce5fcbba-f310-4d20-b828-11e8950fc545Cited by top-tier papers2
- PartialNet: Compute Less, Perform BetterHaiduo Huang, Tian Xia, Wenzhe Zhao, Pengju RenAAAI 2026
- SF-Mamba: Rethinking State Space Model for VisionMasakazu Yoshimura, Teruaki Hayashi, Yuki Hoshino, Wei-Yao Wang et al.ICML 2026
Builds on43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.ICCV 2023 · 341 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
- SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision ApplicationsAbdelrahman M. Shaker, Muhammad Maaz, Hanoona Abdul Rasheed, Salman H. Khan et al.ICCV 2023 · 213 citations
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICLR 2024 · 30 citations
