Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers
Bin Ren, Yahui Liu, Yue Song, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang
摘要
Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the nonoccluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Sharing Key Semantics in Transformer Makes Efficient Image RestorationBin Ren, Yawei Li, Jingyun Liang, Rakesh Ranjan 等NeurIPS 2024 · 被引用 17 次
- Linear Mechanisms for Spatiotemporal Reasoning in Vision Language ModelsRaphaela Kang, Hongqiao Chen, Georgia Gkioxari, Pietro PeronaICLR 2026 · 被引用 11 次
- Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold RegularizationBin Ren, Yawei Li, Xu Zheng, Yuqian Fu 等ICLR 2026 · 被引用 9 次
- ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded GapsXingke Song, Xiaoying Yang, Chenglin Yao, Jianfeng Ren 等AAAI 2025 · 被引用 6 次
- SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language PretrainingYue Li, Qi Ma, Runyi Yang, Huapeng Li 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
相关 Paper
- Membership Inference Attacks against Vision Transformers: Mosaic MixUp Training to the DefenseQiankun Zhang, Di Yuan, Boyu Zhang, Bin Yuan 等CCS 2024 · 被引用 1 次
- Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision TransformersBum Jun Kim, Sang Woo KimAAAI 2025 · 被引用 2 次
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 被引用 115 次
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 被引用 1 次
- When Adversarial Training Meets Vision Transformers: Recipes from Training to ArchitectureYichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo 等NeurIPS 2022 · 被引用 109 次
