PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers
Ryan Grainger, Thomas Paniagua, Xi Song, Naresh P. Cuntoor, Mun Wai Lee, Tianfu Wu
Abstract
Vision Transformers (ViTs) are built on the assumption of treating image patches as "visual tokens" and learn patch-to-patch attention. The patch embedding based tokenizer has a semantic gap with respect to its counterpart, the textual tokenizer. The patch-to-patch attention suffers from the quadratic complexity issue, and also makes it nontrivial to explain learned ViTs. To address these issues in ViT, this paper proposes to learn Patch-to-Cluster attention (PaCa) in ViT. Queries in our PaCa-ViT starts with patches, while keys and values are directly based on clustering (with a predefined small number of clusters). The clusters are learned end-to-end, leading to better tokenizers and inducing joint clustering-for-attention and attentionfor-clustering for better and interpretable models. The quadratic complexity is relaxed to linear complexity. The proposed PaCa module is used in designing efficient and interpretable ViT backbones and semantic segmentation head networks. In experiments, the proposed methods are tested on ImageNet-1k image classification, MS-COCO object detection and instance segmentation and MIT-ADE20k semantic segmentation. Compared with the prior art, it obtains better performance in all the three benchmarks than the SWin [32] and the PVTs [47, 48] by significant margins in ImageNet-1k and MIT-ADE20k. It is also significantly more efficient than PVT models in MS-COCO and MIT-ADE20k due to the linear complexity. The learned clusters are semantically meaningful. Code and model checkpoints are available at https://github.com/ iVMCL/PaCaViT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a8d73d4-9b0d-460b-8fd3-7bccc52c171eCited by top-tier papers4
- Revisiting the Integration of Convolution and Attention for Vision BackboneLei Zhu, Xinjiang Wang, Wayne Zhang, Rynson W. H. LauNeurIPS 2024 · 9 citations
- Multi-View Attentive Contextualization for Multi-View 3D Object DetectionXianpeng Liu, Ce Zheng, Ming Qian, Nan Xue et al.CVPR 2024 · 5 citations
- Dual-Temporal Exemplar Representation Network for Video Semantic SegmentationXiaolong Xu, Lei Zhang, Jiayi Li, Lituan Wang et al.ICCV 2025 · 3 citations
- Efficient Data Driven Mixture-of-Expert Extraction from Trained NetworksUranik Berisha, Jens Mehnert, Alexandru Paul ConduracheCVPR 2025
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
Related papers
- SegViT: Semantic Segmentation with Plain Vision TransformersBowen Zhang, Zhi Tian, Quan Tang, Xiangxiang Chu et al.NeurIPS 2022 · 242 citations
- Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision TokensQihang Fan, Huaibo Huang, Mingrui Chen, Ran HeICCV 2025 · 1 citation
- Making Vision Transformers Efficient from A Token Sparsification ViewShuning Chang, Pichao Wang, Ming Lin, Fan Wang et al.CVPR 2023
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song et al.ICLR 2022 · 27 citations
