Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting
Zhicheng Wang, Liwen Xiao, Zhiguo Cao, Hao Lu
Abstract
Class-agnostic counting (CAC) aims to count objects of interest from a query image given few exemplars. This task is typically addressed by extracting the features of query image and exemplars respectively and then matching their feature similarity, leading to an extract-then-match paradigm. In this work, we show that CAC can be simplified in an extract-and-match manner, particularly using a vision transformer (ViT) where feature extraction and similarity matching are executed simultaneously within the self-attention. We reveal the rationale of such simplification from a decoupled view of the self-attention.The resulting model, termed CACViT, simplifies the CAC pipeline into a single pretrained plain ViT. Further, to compensate the loss of the scale and the order-of-magnitude information due to resizing and normalization in plain ViT, we present two effective strategies for scale and magnitude embedding. Extensive experiments on the FSC147 and the CARPK datasets show that CACViT significantly outperforms state-of-the-art CAC approaches in both effectiveness (23.60% error reduction) and generalization, which suggests CACViT provides a concise and strong baseline for CAC. Code will be available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60dc5d85-3ca5-4d6d-92b6-fa66ffb2bcd1Cited by top-tier papers9
- PBECount: Prompt-Before-Extract Paradigm for Class-Agnostic CountingCanchen Yang, Tianyu Geng, Jian Peng, Chun XuAAAI 2025 · 3 citations
- Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant SpeciesJinyu Xu, Tianqi Hu, Xiaonan Hu, Letian Zhou et al.CVPR 2026 · 2 citations
- CountSE: Soft Exemplar Open-Set Object CountingShuai Liu, Peng Zhang, Shiwei Zhang, Wei KeICCV 2025 · 1 citation
- Generalized-Scale Object Counting with Gradual Query AggregationJer Pelhan, Alan Lukezic, Matej KristanAAAI 2026
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
Builds on13
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen et al.CVPR 2022 · 183 citations
- Localization in the Crowd with Topological ConstraintsShahira Abousamra, Minh Hoai, Dimitris Samaras, Chao ChenAAAI 2021 · 160 citations
Related papers
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu et al.CVPR 2022 · 99 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
