GhostNetV2: Enhance Cheap Operation with Long-Range Attention
Yehui Tang, Kai Han, Jianyuan Guo, Chang Xu, Chao Xu, Yunhe Wang
Abstract
Light-weight convolutional neural networks (CNNs) are specially designed for applications on mobile devices with faster inference speed. The convolutional operation can only capture local information in a window region, which prevents performance from being further improved. Introducing self-attention into convolution can capture global information well, but it will largely encumber the actual speed. In this paper, we propose a hardware-friendly attention mechanism (dubbed DFC attention) and then present a new GhostNetV2 architecture for mobile applications. The proposed DFC attention is constructed based on fully-connected layers, which can not only execute fast on common hardware but also capture the dependence between long-range pixels. We further revisit the expressiveness bottleneck in previous GhostNet and propose to enhance expanded features produced by cheap operations with DFC attention, so that a GhostNetV2 block can aggregate local and long-range information simultaneously. Extensive experiments demonstrate the superiority of GhostNetV2 over existing architectures. For example, it achieves 75.3% top-1 accuracy on ImageNet with 167M FLOPs, significantly suppressing GhostNetV1 (74.5%) with a similar computational cost. The source code will be available at https://github.com/huawei-noah/Efficient-AI-Backbones/tree/master/ghostnetv2_pytorch and https://gitee.com/mindspore/models/tree/master/research/cv/ghostnetv2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a8f05b3-249c-4fd3-91b5-5fa061233fd4Cited by top-tier papers10
- Efficient Modulation for Vision NetworksXu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICLR 2024 · 30 citations
- LWGANet: Addressing Spatial and Channel Redundancy in Remote Sensing Visual Tasks with Light-Weight Grouped AttentionWei Lu, Xue Yang, Si-Bao ChenAAAI 2026 · 21 citations
- MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile DevicesHailong Yan, Ao Li, Xiangtao Zhang, Zhe Liu et al.ICCV 2025 · 12 citations
- Towards Real-time Video Compressive Sensing on Mobile DevicesMiao Cao, Lishun Wang, Huan Wang, Guoqing Wang et al.ACM MM 2024 · 4 citations
- COSMIC: Compress Satellite Image Efficiently via Diffusion CompensationZiyuan Zhang, Han Qiu, Maosen Zhang, Jun Liu et al.NeurIPS 2024 · 4 citations
Builds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
Related papers
- GhostNet: More Features From Cheap OperationsKai Han, Yunhe Wang, Qi Tian, Jianyuan Guo et al.CVPR 2020
- Neural Architecture Search for Lightweight Non-Local NetworksYingwei Li, Xiaojie Jin, Jieru Mei, Xiaochen Lian et al.CVPR 2020
- Dynamic Convolution: Attention Over Convolution KernelsYinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen et al.CVPR 2020
- Iformer: Integrating ConvNet and Transformer for Mobile ApplicationChuanyang ZhengICLR 2025
- Coordinate Attention for Efficient Mobile Network DesignQibin Hou, Daquan Zhou, Jiashi FengCVPR 2021
