Keep It SimPool: Who Said Supervised Transformers Suffer from Attention Deficit?
Bill Psomas, Ioannis Kakogeorgiou, Konstantinos Karantzalos, Yannis Avrithis
Abstract
Convolutional networks and vision transformers have different forms of pairwise interactions, pooling across layers and pooling at the end of the network. Does the latter really need to be different? As a by-product of pooling, vision transformers provide spatial attention for free, but this is most often of low quality unless self-supervised, which is not well studied. Is supervision really the problem?In this work, we develop a generic pooling framework and then we formulate a number of existing methods as instantiations. By discussing the properties of each group of methods, we derive SimPool, a simple attention-based pooling mechanism as a replacement of the default one for both convolutional and transformer encoders. We find that, whether supervised or self-supervised, this improves performance on pre-training and downstream tasks and provides attention maps delineating object boundaries in all cases. One could thus call SimPool universal. To our knowledge, we are the first to obtain attention maps in supervised transformers of at least as good quality as self-supervised, without explicit losses or modifying the architecture. Code at: https://github.com/billpsomas/simpool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 981fe6ca-e045-4435-bfc0-cb75a461b192Cited by top-tier papers9
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 17 citations
- Attention, Please! Revisiting Attentive Probing Through the Lens of EfficiencyBill Psomas, Dionysis Christopoulos, Eirini Baltzi, Ioannis Kakogeorgiou et al.ICLR 2026 · 12 citations
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTsAlexander Lappe, Martin A. GieseNeurIPS 2025 · 9 citations
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio ClassificationLukas Rauch, René Heinrich, Houtan Ghaffari, Lukas Miklautz et al.ICLR 2026 · 7 citations
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Efficient Representation Learning via Adaptive Context PoolingChen Huang, Walter Talbott, Navdeep Jaitly, Joshua M. SusskindICML 2022 · 10 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsSofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang et al.NeurIPS 2025 · 6 citations
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 35 citations
- Spatially Attentive Output Layer for Image ClassificationIldoo Kim, Woonhyuk Baek, Sungwoong KimCVPR 2020
