Global Context Vision Transformers
Ali Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz, Pavlo Molchanov
Abstract
We propose global context vision transformer (GC ViT), a novel architecture that enhances parameter and compute utilization for computer vision. Our method leverages global context self-attention modules, joint with standard local self-attention, to effectively and efficiently model both long and short-range spatial interactions, without the need for expensive operations such as computing attention masks or shifting local windows. In addition, we address the lack of the inductive bias in ViTs, and propose to leverage a modified fused inverted residual blocks in our architecture. Our proposed GC ViT achieves state-of-the-art results across image classification, object detection and semantic segmentation tasks. On ImageNet-1K dataset for classification, the variants of GC ViT with 51M, 90M and 201M parameters achieve 84.3%, 85.0% and 85.7% Top-1 accuracy, respectively, at 224 × 224 image resolution and without any pre-training, hence surpassing comparably-sized prior art such as CNN-based ConvNeXt and ViTbased MaxViT and Swin Transformer by a large margin. Pre-trained GC ViT backbones in downstream tasks of object detection, instance segmentation, and semantic segmentation using MS COCO and ADE20K datasets outperform prior work consistently. Specifically, GC ViT with a 4scale DINO detection head achieves a box AP of 58.3% on MS COCO dataset. Code is available at https://github.com/NVlabs/GCViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ab99afb-f0ee-40ed-9c6c-15a523beffa4Cited by top-tier papers34
- U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationChenxin Li, Xinyu Liu, Wuyang Li, Cheng Wang et al.AAAI 2025 · 452 citations
- Locality-Attending Vision TransformerSina Hajimiri, Farzad Beizaee, Fereshteh Shakeri, Christian Desrosiers et al.ICLR 2026 · 427 citations
- FasterViT: Fast Vision Transformers with Hierarchical AttentionAli Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao et al.ICLR 2024 · 132 citations
- FFT-Based Dynamic Token Mixer for VisionYuki Tatsunami, Masato TakiAAAI 2024 · 73 citations
- Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision MakingJeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul SungICLR 2024 · 36 citations
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual RecognitionHao Liu, Xinghua Jiang, Xin Li, Zhimin Bao et al.CVPR 2022 · 17 citations
- GPViT: A High Resolution Non-Hierarchical Vision Transformer with Group PropagationChenhongyi Yang, Jiarui Xu, Shalini De Mello, Elliot J. Crowley et al.ICLR 2023 · 7 citations
- Pale Transformer: A General Vision Transformer Backbone with Pale-Shaped AttentionSitong Wu, Tianyi Wu, Haoru Tan, Guodong GuoAAAI 2022 · 84 citations
- Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionYuxin Fang, Shusheng Yang, Shijie Wang, Yixiao Ge et al.ICCV 2023 · 67 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
