Glance-and-Gaze Vision Transformer
Qihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu, Alan L. Yuille, Wei Shen
Abstract
Recently, there emerges a series of vision Transformers, which show superior performance with a more compact model size than conventional convolutional neural networks, thanks to the strong ability of Transformers to model long-range dependencies. However, the advantages of vision Transformers also come with a price: Self-attention, the core part of Transformer, has a quadratic complexity to the input sequence length. This leads to a dramatic increase of computation and memory cost with the increase of sequence length, thus introducing difficulties when applying Transformers to the vision tasks that require dense predictions based on high-resolution feature maps. In this paper, we propose a new vision Transformer, named Glance-and-Gaze Transformer (GG-Transformer), to address the aforementioned issues. It is motivated by the Glance and Gaze behavior of human beings when recognizing objects in natural scenes, with the ability to efficiently model both long-range dependencies and local context. In GG-Transformer, the Glance and Gaze behavior is realized by two parallel branches: The Glance branch is achieved by performing self-attention on the adaptively-dilated partitions of the input, which leads to a linear complexity while still enjoying a global receptive field; The Gaze branch is implemented by a simple depth-wise convolutional layer, which compensates local image context to the features obtained by the Glance mechanism. We empirically demonstrate our method achieves consistently superior performance over previous state-of-the-art Transformers on various vision tasks and benchmarks. The codes and models will be made available at https://github.com/yucornetto/GG-Transformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 2,652 citations
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 339 citations
- CF-ViT: A General Coarse-to-Fine Method for Vision TransformerMengzhao Chen, Mingbao Lin, Ke Li, Yunhang Shen et al.AAAI 2023 · 105 citations
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger TokensJiemin Fang, Lingxi Xie, Xinggang Wang, Xiaopeng Zhang et al.CVPR 2022 · 73 citations
- Scaled ReLU Matters for Training Vision TransformersPichao Wang, Xue Wang, Hao Luo, Jingkai Zhou et al.AAAI 2022 · 55 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai et al.NeurIPS 2021 · 228 citations
- Less Is More: Pay Less Attention in Vision TransformersZizheng Pan, Bohan Zhuang, Haoyu He, Jing Liu et al.AAAI 2022 · 109 citations
- Group Vision TransformerYaopeng Peng, Milan Sonka, Danny Z. ChenACM MM 2024
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image EncodingPengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICCV 2021 · 384 citations
- Learned Queries for Efficient Local AttentionMoab Arar, Ariel Shamir, Amit H. BermanoCVPR 2022 · 28 citations
