Permutation Equivariance of Transformers and its Applications
Hengyuan Xu, Liyao Xiang, Hangyu Ye, Dixi Yao, Pengzhi Chu, Baochun Li
Abstract
Revolutionizing the field of deep learning, Transformer-based models have achieved remarkable performance in many tasks. Recent research has recognized these models are robust to shuffling but are limited to inter-token permutation in the forward propagation. In this work, we propose our definition of permutation equivariance, a broader concept covering both inter- and intra- token per-mutation in the forward and backward propagation of neural networks. We rigorously proved that such permutation equivariance property can be satisfied on most vanilla Transformer-based models with almost no adaptation. We examine the property over a range of state-of-the-art models including ViT, Bert, GPT, and others, with experimental validations. Further, as a proof-of-concept, we explore how real-world applications including privacy-enhancing split learning, and model authorization, could exploit the permutation equivariance property, which implicates wider, intriguing application scenarios. The code is available at https://github.com/Doby-Xu/ST
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3581c781-4824-4b27-8a97-0819640930e8Cited by top-tier papers11
- WithAnyone: Toward Controllable and ID Consistent Image GenerationHengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang et al.ICLR 2026 · 12 citations
- SPINT: Spatial Permutation-Invariant Neural Transformer for Consistent Intracortical Motor DecodingTrung Le, Hao Fang, Jingyuan Li, Tung Nguyen et al.NeurIPS 2025 · 8 citations
- Maximizing the Position Embedding for Vision Transformers with Global Average PoolingWonjun Lee, Bumsub Ham, Suhyun KimAAAI 2025 · 3 citations
- REOrdering Patches Improves Vision ModelsDeclan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell et al.NeurIPS 2025 · 3 citations
- STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in ProductionMu Yuan, Lan Zhang, Yihang Cheng, Miao-Hui Song et al.NDSS 2026 · 2 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
Related papers
- Split Adaptation for Pre-trained Vision TransformersLixu Wang, Bingqi Shang, Yi Li, Payal Mohapatra et al.CVPR 2025
- How Does a Deep Learning Model Architecture Impact Its Privacy? A Comprehensive Study of Privacy Attacks on CNNs and TransformersGuangsheng Zhang, Bo Liu, Huan Tian, Tianqing Zhu et al.USENIX Security 2024
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- In What Ways Are Deep Neural Networks Invariant and How Should We Measure This?Henry Kvinge, Tegan Emerson, Grayson Jorgenson, Scott Vasquez et al.NeurIPS 2022 · 15 citations
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 1 citation
