Content-aware Token Sharing for Efficient Semantic Segmentation with Vision Transformers
Chenyang Lu, Daan de Geus, Gijs Dubbelman
Abstract
This paper introduces Content-aware Token Sharing (CTS), a token reduction approach that improves the computational efficiency of semantic segmentation networks that use Vision Transformers (ViTs). Existing works have proposed token reduction approaches to improve the efficiency of ViT-based image classification networks, but these methods are not directly applicable to semantic segmentation, which we address in this work. We observe that, for semantic segmentation, multiple image patches can share a token if they contain the same semantic class, as they contain redundant information. Our approach leverages this by employing an efficient, class-agnostic policy network that predicts if image patches contain the same semantic class, and lets them share a token if they do. With experiments, we explore the critical design choices of CTS and show its effectiveness on the ADE20K, Pascal Context and Cityscapes datasets, various ViT backbones, and different segmentation decoders. With Content-aware Token Sharing, we are able to reduce the number of processed tokens by up to 44%, without diminishing the segmentation quality.
- Both authors contributed equally. 1 In this work we use the term ViTs for the complete family of vision transformers that purely apply global self-attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a6f0398-457b-4197-98e7-ac3df06c16dbCited by top-tier papers10
- DiTFastAttn: Attention Compression for Diffusion Transformer ModelsZhihang Yuan, Hanling Zhang, Lu Pu, Xuefei Ning et al.NeurIPS 2024 · 134 citations
- Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationQuan Tang, Bowen Zhang, Jiajun Liu, Fagui Liu et al.ICCV 2023 · 74 citations
- SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationBin Xie, Jiale Cao, Jin Xie, Fahad Shahbaz Khan et al.CVPR 2024 · 57 citations
- A Closer Look at the CLS Token for Cross-Domain Few-Shot LearningYixiong Zou, Shuai Yi, Yuhua Li, Ruixuan LiNeurIPS 2024 · 40 citations
- TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing ModalitiesZheyu Zhang, Gang Yang, Yueyi Zhang, Huanjing Yue et al.AAAI 2024 · 29 citations
Builds on32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- Making Vision Transformers Efficient from A Token Sparsification ViewShuning Chang, Pichao Wang, Ming Lin, Fan Wang et al.CVPR 2023
- ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision TransformersNarges Norouzi, Svetlana Orlova, Daan de Geus, Gijs DubbelmanCVPR 2024 · 15 citations
- SegViT: Semantic Segmentation with Plain Vision TransformersBowen Zhang, Zhi Tian, Quan Tang, Xiangxiang Chu et al.NeurIPS 2022 · 242 citations
- PaCa-ViT: Learning Patch-to-Cluster Attention in Vision TransformersRyan Grainger, Thomas Paniagua, Xi Song, Naresh P. Cuntoor et al.CVPR 2023
