Scaled ReLU Matters for Training Vision Transformers
Pichao Wang, Xue Wang, Hao Luo, Jingkai Zhou, Zhipeng Zhou, Fan Wang, Hao Li, Rong Jin
Abstract
Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in (Xiao et al. 2021) , and the authors conjecture that the issue lies with the patchify-stem of ViT models and propose that early convolutions help transformers see better. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11864813-a2b2-40e7-bbb8-a2cbd46388e6Cited by top-tier papers9
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Shunted Self-Attention via Multi-Scale Token AggregationSucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng et al.CVPR 2022 · 326 citations
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou et al.ICLR 2022 · 168 citations
- VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsZhenyu Wang, Hao Luo, Pichao Wang, Feng Ding et al.NeurIPS 2022 · 57 citations
- Choose Wisely: An Extensive Evaluation of Model Selection for Anomaly Detection in Time SeriesEmmanouil Sylligardos, Paul Boniol, John Paparrizos, Panos E. Trahanias et al.VLDB 2023 · 40 citations
Builds on28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song et al.ICLR 2022 · 27 citations
- A General and Efficient Training for Transformer via Token ExpansionWenxuan Huang, Yunhang Shen, Jiao Xie, Baochang Zhang et al.CVPR 2024
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah et al.CVPR 2022 · 36 citations
- Bootstrapping ViTs: Towards Liberating Vision Transformers from Pre-trainingHaofei Zhang, Jiarui Duan, Mengqi Xue, Jie Song et al.CVPR 2022 · 19 citations
