Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers
Bin Ren, Yahui Liu, Yue Song, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang
Abstract
Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the nonoccluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d7d0af5-17bf-435a-995c-663a70aa6c8dCited by top-tier papers11
- Sharing Key Semantics in Transformer Makes Efficient Image RestorationBin Ren, Yawei Li, Jingyun Liang, Rakesh Ranjan et al.NeurIPS 2024 · 17 citations
- Linear Mechanisms for Spatiotemporal Reasoning in Vision Language ModelsRaphaela Kang, Hongqiao Chen, Georgia Gkioxari, Pietro PeronaICLR 2026 · 11 citations
- Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold RegularizationBin Ren, Yawei Li, Xu Zheng, Yuqian Fu et al.ICLR 2026 · 9 citations
- ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded GapsXingke Song, Xiaoying Yang, Chenglin Yao, Jianfeng Ren et al.AAAI 2025 · 6 citations
- SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language PretrainingYue Li, Qi Ma, Runyi Yang, Huapeng Li et al.ICCV 2025 · 5 citations
Builds on28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
Related papers
- Membership Inference Attacks against Vision Transformers: Mosaic MixUp Training to the DefenseQiankun Zhang, Di Yuan, Boyu Zhang, Bin Yuan et al.CCS 2024 · 1 citation
- Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision TransformersBum Jun Kim, Sang Woo KimAAAI 2025 · 2 citations
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 115 citations
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 1 citation
- When Adversarial Training Meets Vision Transformers: Recipes from Training to ArchitectureYichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo et al.NeurIPS 2022 · 109 citations
