Dynamic Token Normalization improves Vision Transformers
Wenqi Shao, Yixiao Ge, Zhaoyang Zhang, Xuyuan Xu, Xiaogang Wang, Ying Shan, Ping Luo
Abstract
Vision Transformer (ViT) and its variants (e.g., Swin, PVT) have achieved great success in various computer vision tasks, owing to their capability to learn long-range contextual information. Layer Normalization (LN) is an essential ingredient in these models. However, we found that the ordinary LN makes tokens at different positions similar in magnitude because it normalizes embeddings within each token. It is difficult for Transformers to capture inductive bias such as the positional context in an image with LN. We tackle this problem by proposing a new normalizer, termed Dynamic Token Normalization (DTN), where normalization is performed both within each token (intra-token) and across different tokens (inter-token). DTN has several merits. Firstly, it is built on a unified formulation and thus can represent various existing normalization methods. Secondly, DTN learns to normalize tokens in both intra-token and inter-token manners, enabling Transformers to capture both the global contextual information and the local positional context. Thirdly, by simply replacing LN layers, DTN can be readily plugged into various vision transformers, such as ViT, Swin, PVT, LeViT, T2T-ViT, BigBird and Reformer. Extensive experiments show that the transformer equipped with DTN consistently outperforms baseline model with minimal extra parameters and computational overhead. For example, DTN outperforms LN by - top-1 accuracy on ImageNet, by - box AP in object detection on COCO benchmark, by - mCE in robustness experiments on ImageNet-C, and by - accuracy in Long ListOps on Long-Range Arena. Codes will be made public at https://github.com/wqshao126/DTN
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cac101f-b853-4d71-976e-c01244df6d35Cited by top-tier papers3
- Cached Transformers: Improving Transformers with Differentiable Memory CachdeZhaoyang Zhang, Wenqi Shao, Yixiao Ge, Xiaogang Wang et al.AAAI 2024 · 7 citations
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video GenerationHaoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang et al.CVPR 2026 · 5 citations
- Unified Normalization for Accelerating and Stabilizing TransformersQiming Yang, Kai Zhang, Chaoxiang Lan, Zhi Yang et al.ACM MM 2022 · 1 citation
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer NormalizationRunyi Yu, Zhennan Wang, Yinhuai Wang, Kehan Li et al.ICCV 2023 · 13 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Beyond Fixation: Dynamic Window Visual TransformerPengzhen Ren, Changlin Li, Guangrun Wang, Yun Xiao et al.CVPR 2022 · 43 citations
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang et al.NeurIPS 2021 · 283 citations
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 1 citation
