Scale-space Tokenization for Improving the Robustness of Vision Transformers
Lei Xu, Rei Kawakami, Nakamasa Inoue
Abstract
The performance of the Vision Transformer (ViT) model and its variants in most vision tasks has surpassed traditional Convolutional Neural Networks (CNNs) in terms of in-distribution accuracy. However, ViTs still have significant room for improvement in their robustness to input perturbations. Furthermore, robustness is a critical aspect to consider when deploying ViTs in real-world scenarios. Despite this, some variants of ViT improve the in-distribution accuracy and computation performance at the cost of sacrificing the model's robustness and generalization. In this study, inspired by the prior findings on the potential effectiveness of shape bias to robustness improvement and the importance of multi-scale analysis, we propose a simple yet effective method, scale-space tokenization, to improve the robustness of ViT while maintaining in-distribution accuracy. Based on this method, we build Scale-space-based Robust Vision Transformer (SRVT) model. Our method consists of scale-space patch embedding and scale-space positional encoding. The scale-space patch embedding makes a sequence of variable-scale images and increases the model's shape bias to enhance its robustness. The scale-space positional encoding implicitly boosts the model's invariance to input perturbations by incorporating scale-aware position information into 3D sinusoidal positional encoding. We conduct experiments on image recognition benchmarks (CIFAR10/100 and ImageNet-1k) from the perspectives of in-distribution accuracy, adversarial and out-of-distribution robustness. The experimental results demonstrate our method's effectiveness in improving robustness without compromising in-distribution accuracy. Especially, our approach achieves advanced adversarial robustness on ImageNet-1k benchmark compared with state-of-the-art robust ViT.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ca9ff6c1-7db7-4c09-bf34-42d2df0c08edRelated papers
- Towards Robust Vision TransformerXiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li et al.CVPR 2022 · 185 citations
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick et al.ICLR 2022 · 47 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision TransformersNick Nikzad, Yi Liao, Yongsheng Gao, Jun ZhouCVPR 2025
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
