Scaling Local Self-Attention for Parameter Efficient Visual Backbones
Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake A. Hechtman, Jonathon Shlens
Abstract
Self-attention has the promise of improving computer vision systems due to parameter-independent scaling of receptive fields and content-dependent interactions, in contrast to parameter-dependent scaling and content-independent interactions of convolutions. Self-attention models have recently been shown to have encouraging improvements on accuracy-parameter trade-offs compared to baseline convolutional models such as ResNet-50. In this work, we aim to develop self-attention models that can outperform not just the canonical baseline models, but even the high-performing convolutional models. We propose two extensions to selfattention that, in conjunction with a more efficient implementation of self-attention, improve the speed, memory usage, and accuracy of these models. We leverage these improvements to develop a new self-attention model family, HaloNets, which reach state-of-the-art accuracies on the parameterlimited setting of the ImageNet classification benchmark. In preliminary transfer learning experiments, we find that HaloNet models outperform much larger models and have better inference performance. On harder tasks such as object detection and instance segmentation, our simple local self-attention and convolutional hybrids show improvements over very strong baselines. These results mark another step in demonstrating the efficacy of self-attention models on settings traditionally dominated by convolutional models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers97
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee et al.SIGGRAPH 2022 · 1,638 citations
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 1,298 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
Related papers
- Bottleneck Transformers for Visual RecognitionAravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens et al.CVPR 2021
- LambdaNetworks: Modeling long-range Interactions without AttentionIrwan BelloICLR 2021 · 48 citations
- Explicitly Modeled Attention Maps for Image ClassificationAndong Tan, Duc Tam Nguyen, Maximilian Dax, Matthias Nießner et al.AAAI 2021 · 10 citations
- Iformer: Integrating ConvNet and Transformer for Mobile ApplicationChuanyang ZhengICLR 2025
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai et al.NeurIPS 2021 · 228 citations
