PBECount: Prompt-Before-Extract Paradigm for Class-Agnostic Counting
Canchen Yang, Tianyu Geng, Jian Peng, Chun Xu
Abstract
In the field of class-agnostic counting (CAC), counting only objects of interest that are similar to exemplars in multi-class scenarios has been a challenging task. To address this challenge, recent research has proposed the extract-and-match paradigm based on the vision transformer (ViT) architecture. However, although this paradigm can improve the accuracy of exemplar-similar object identification, it overly emphasizes the role of the ViT structure. To address this shortcoming, this work introduces a more generalized prompt-before-extract paradigm on top of the extract-and-match paradigm and designs a pure convolutional neural network (CNN) model named PBECount. In addition, an innovative loss function, a post-processing strategy, and a dynamic threshold method are proposed to enhance the detection performance of the proposed model when the probability maps are used as ground truth during model training. The experimental results on the FSC-147 and CARPK datasets demonstrate that the proposed PBECount can identify whether unknown class objects are similar to exemplars and outperform the state-of-the-art CAC methods in terms of accuracy and generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu et al.CVPR 2022 · 99 citations
- Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic CountingZhicheng Wang, Liwen Xiao, Zhiguo Cao, Hao LuAAAI 2024 · 35 citations
- DAVE - A Detect-and-Verify Paradigm for Low-Shot CountingJer Pelhan, Alan Lukezic, Vitjan Zavrtanik, Matej KristanCVPR 2024 · 14 citations
- Learning To Count EverythingViresh Ranjan, Udbhav Sharma, Thu Nguyen, Minh HoaiCVPR 2021
Related papers
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- A Fixed-Point Approach to Unified Prompt-Based CountingWei Lin, Antoni B. ChanAAAI 2024 · 11 citations
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang et al.CVPR 2022 · 14 citations
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
- Guided Attention Network for Object Detection and Counting on DronesYuanqiang Cai, Dawei Du, Libo Zhang, Longyin Wen et al.ACM MM 2020 · 60 citations
