Making Vision Transformers Truly Shift-Equivariant
Renan A. Rojas-Gomez, Teck-Yian Lim, Minh N. Do, Raymond A. Yeh
Abstract
In the field of computer vision, Vision Transformers (ViTs) have emerged as a prominent deep learning architecture. Despite being inspired by Convolutional Neural Networks (CNNs), ViTs are susceptible to small spatial shifts in the input data - they lack shift-equivariance. To address this shortcoming, we introduce novel data-adaptive designs for each of the ViT modules that break shift-equivariance, such as tokenization. self-attention, patch merging, and positional encoding. With our proposed modules, we achieve perfect circular shift-equivariance across four prominent ViT ar-chitectures: Swin, SwinV2, CvT, and MViTv2. Additionally, we leverage our design to further enhance consistency under standard shifts. We evaluate our adaptive ViT models on image classification and semantic segmentation tasks. Our models achieve competitive performance across three diverse datasets, showcasing perfect (100%) circular shift consistency while improving standard shift consistency. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Project website: https://renanrojasg.github.io/shifteq_vit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df8a00f2-8988-4bbd-bbe3-eadc0a3ce44aCited by top-tier papers9
- Truly Scale-Equivariant Deep Nets with Fourier LayersMd Ashiqur Rahman, Raymond A. YehNeurIPS 2023 · 17 citations
- Enhancing Image Restoration Transformer via Adaptive Translation EquivarianceJiaKui Hu, Zhengjian Yao, Lujia Jin, Hangzhou He et al.ICCV 2025 · 6 citations
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness AsymmetryXinchang Wang, Yunhao Chen, Yuechen Zhang, Congcong Bian et al.ICML 2026 · 2 citations
- CLIPSym: Delving into Symmetry Detection with CLIPTinghan Yang, Md Ashiqur Rahman, Raymond A. YehICCV 2025 · 2 citations
- Alias-Free ViT: Fractional Shift Invariance via Linear AttentionHagay Michaeli, Daniel SoudryNeurIPS 2025 · 2 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Vision Transformers Are Robust LearnersSayak Paul, Pin-Yu ChenAAAI 2022 · 372 citations
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- Towards Robust Vision TransformerXiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li et al.CVPR 2022 · 185 citations
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 1 citation
