Focal Attention for Long-Range Interactions in Vision Transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, Jianfeng Gao
Abstract
Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing local and global visual dependencies through self-attention is the key to its success. However, this also brings challenges due to quadratic computational overhead, especially for the high-resolution vision tasks (e.g., object detection). Many recent works have attempted to reduce the cost and improve model performance by applying either coarse-grained global attention or fine-grained local attention. However, both approaches cripple the modeling power of the original self-attention mechanism of multi-layer Transformers, leading to sub-optimal solutions. In this paper, we present focal attention, a new attention mechanism that incorporates both fine-grained local and coarse-grained global interactions. In this new mechanism, each token attends its closest surrounding tokens at fine granularity and the tokens far away at coarse granularity, and thus can capture both short-and long-range visual dependencies efficiently and effectively. With focal attention, we build a new variant of Vision Transformer models, called Focal Transformers, which achieve superior performance over the state-of-theart (SoTA) Vision Transformers on a range of public image classification and object detection benchmarks. In particular, our Focal Transformer models with a moderate size of 51.1M and a large size of 89.8M achieve 83.6% and 84.0% Top-1 accuracy, respectively, on ImageNet classification at 224 × 224. When employed as the backbones, Focal Transformers achieve consistent and substantial improvements over the current SoTA Swin Transformers [43] across 6 different object detection methods. Our largest Focal Transformer yields 58.7/59.0 box mAPs and 50.9/51.3 mask mAPs on COCO mini-val/test-dev, and 55.4 mIoU on ADE20K for semantic segmentation, creating new SoTA on three of the most challenging computer vision tasks. Our code is available at: https://github. com/microsoft/Focal-Transformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ffe2941-eb2e-4250-8e36-b61b3a64431fCited by top-tier papers38
- HorNet: Efficient High-Order Spatial Interactions with Recursive Gated ConvolutionsYongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou et al.NeurIPS 2022 · 422 citations
- Rethinking Mobile Block for Efficient Attention-based ModelsJiangning Zhang, Xiangtai Li, Jian Li, Liang Liu et al.ICCV 2023 · 223 citations
- Global Context Vision TransformersAli Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz et al.ICML 2023 · 213 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- Unified Contrastive Learning in Image-Text-Label SpaceJianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao et al.CVPR 2022 · 182 citations
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 494 citations
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 70 citations
- Pale Transformer: A General Vision Transformer Backbone with Pale-Shaped AttentionSitong Wu, Tianyi Wu, Haoru Tan, Guodong GuoAAAI 2022 · 84 citations
- Group Vision TransformerYaopeng Peng, Milan Sonka, Danny Z. ChenACM MM 2024
- Lightweight Vision Transformer with Bidirectional InteractionQihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran HeNeurIPS 2023 · 59 citations
