Point, Segment and Count: A Generalized Framework for Object Counting
Zhizhong Huang, Mingliang Dai, Yi Zhang, Junping Zhang, Hongming Shan
Abstract
Class-agnostic object counting aims to count all objects in an image with respect to example boxes or class names, a.k.a few-shot and zero-shot counting. In this paper, we propose a generalized framework for both few-shot and zeroshot object counting based on detection. Our framework combines the superior advantages of two foundation models without compromising their zero-shot capability: (i) SAM to segment all possible objects as mask proposals, and (ii) CLIP to classify proposals to obtain accurate object counts. However, this strategy meets the obstacles of efficiency overhead and the small crowded objects that cannot be localized and distinguished. To address these issues, our framework, termed PseCo, follows three steps: point, segment, and count. Specifically, we first propose a class-agnostic object localization to provide accurate but least point prompts for SAM, which consequently not only reduces computation costs but also avoids missing small objects. Furthermore, we propose a generalized object classification that leverages CLIP image/text embeddings as the classifier, following a hierarchical knowledge distillation to obtain discriminative classifications among hierarchical mask proposals. Extensive experimental results on FSC-147, COCO, and LVIS demonstrate that PseCo achieves state-of-the-art performance in both few-shot/zero-shot object counting/detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- A Novel Unified Architecture for Low-Shot Counting by Detection and SegmentationJer Pelhan, Alan Lukezic, Vitjan Zavrtanik, Matej KristanNeurIPS 2024 · 29 citations
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 14 citations
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao et al.CVPR 2026 · 6 citations
- Fully Exploiting Every Real Sample: SuperPixel Sample Gradient Model StealingYunlong Zhao, Xiaoheng Deng, Yijing Liu, Xinjun Pei et al.CVPR 2024 · 6 citations
- Open-World Object Counting in VideosNiki Amini-Naieni, Andrew ZissermanAAAI 2026 · 6 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
Related papers
- Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIPSepideh Esmaeilpour, Bing Liu, Eric Robertson, Lei ShuAAAI 2022 · 219 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic SegmentationZiqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu et al.CVPR 2023
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- Text Augmented Correlation Transformer For Few-shot Classification & SegmentationSrinivasa Rao Nandam, Sara Atito, Zhenhua Feng, Josef Kittler et al.CVPR 2025
