Mobile Attention: Mobile-Friendly Linear-Attention for Vision Transformers
Zhiyu Yao, Jian Wang, Haixu Wu, Jingdong Wang, Mingsheng Long
摘要
Vision Transformers (ViTs) excel in computer vision tasks due to their ability to capture global context among tokens. However, their quadratic complexity O(N 2 D) in terms of token number N and feature dimension D limits practical use on mobile devices, necessitating more mobilefriendly ViTs with reduced latency. Multi-head linear-attention is emerging as a promising alternative with linear complexity O(N Dd), where d is the per-head dimension. Still, more compute is needed as d gets large for model accuracy. Reducing d improves mobile friendliness at the expense of excessive small heads weak at learning valuable subspaces, ultimately impeding model capability. To overcome this efficiency-capability dilemma, we propose a novel Mobile-Attention design with a head-competition mechanism empowered by information flow, which prevents overemphasis on less important subspaces upon trivial heads while preserving essential subspaces to ensure Transformer's capability. It enables linear-time complexity on mobile devices by supporting a small per-head dimension d for mobile efficiency. By replacing the standard attention of ViTs with Mobile-Attention, our optimized ViTs achieved enhanced model capacity and competitive performance in a range of computer vision tasks. Specifically, we have achieved remarkable reductions in latency on the iPhone 12. Code is available at https://github.com/thuml/MobileAttention .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- PolaFormer: Polarity-aware Linear Attention for Vision TransformersWeikang Meng, Yadan Luo, Xin Li, Dongmei Jiang 等ICLR 2025
- CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual InteractionYuan Zhou, Qingshan Xu, Jiequan Cui, Junbao Zhou 等CVPR 2025
它引用的顶会 Paper27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis 等ICCV 2023 · 被引用 300 次
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han 等CVPR 2024 · 被引用 500 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- You Only Need Less Attention at Each Stage in Vision TransformersShuoxi Zhang, Hanpeng Liu, Stephen Lin, Kun HeCVPR 2024 · 被引用 19 次
