Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
Wonjun Lee, Bumsub Ham, Suhyun Kim
Abstract
In vision transformers, position embedding (PE) plays a crucial role in capturing the order of tokens. However, in vision transformer structures, there is a limitation in the expressiveness of PE due to the structure where position embedding is simply added to the token embedding. A layer-wise method that delivers PE to each layer and applies independent Layer Normalizations for token embedding and PE has been adopted to overcome this limitation. In this paper, we identify the conflicting result that occurs in a layer-wise structure when using the global average pooling (GAP) method instead of the class token. To overcome this problem, we propose MPVG, which maximizes the effectiveness of PE in a layer-wise structure with GAP. Specifically, we identify that PE counterbalances token embedding values at each layer in a layer-wise structure. Furthermore, we recognize that the counterbalancing role of PE is insufficient in the layer-wise structure, and we address this by maximizing the effectiveness of PE through MPVG. Through experiments, we demonstrate that PE performs a counterbalancing role and that maintaining this counterbalancing directionality significantly impacts vision transformers. As a result, the experimental results show that MPVG outperforms existing methods across vision transformers on various tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
Related papers
- LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer NormalizationRunyi Yu, Zhennan Wang, Yinhuai Wang, Kehan Li et al.ICCV 2023 · 13 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
- An Anchor-based Relative Position Embedding Method for Cross-Modal TasksYa Wang, Xingwu Sun, Fengzong Lian, Zhanhui Kang et al.EMNLP 2022 · 1 citation
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang et al.ICML 2026
- GTA: A Geometry-Aware Attention Mechanism for Multi-View TransformersTakeru Miyato, Bernhard Jaeger, Max Welling, Andreas GeigerICLR 2024 · 51 citations
