Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global Attention
Shiwei Zhang, Qi Zhou, Wei Ke
Abstract
Text-guided zero-shot object counting leverages visionlanguage models (VLMs) to count objects of an arbitrary class given by a text prompt. Existing approaches for this challenging task only utilize local patch-level features to fuse with text feature, ignoring the important influence of the global image-level feature. In this paper, we propose a universal strategy that can exploit both local patchlevel features and global image-level feature simultaneously. Specifically, to improve the localization ability of VLMs, we propose Text-guided Local Ranking. Depending on the prior knowledge that foreground patches have higher similarity with the text prompt, a new local-text rank loss is designed to increase the differences between the similarity scores of foreground and background patches which push foreground and background patches apart. To enhance the counting ability of VLMs, Number-evoked Global Attention is introduced to first align global image-level feature with multiple number-conditioned text prompts. Then, the one with the highest similarity is selected to compute cross-attention with the global image-level feature. Through extensive experiments on widely used datasets and methods, the proposed approach has demonstrated superior advancements in performance, generalization, and scalability. Furthermore, to better evaluate text-guided zeroshot object counting methods, we propose a dataset named ZSC-8K, which is larger and more challenging, to establish a new benchmark. Codes and dataset are released at https://github.com/zaqai/LGCount.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8e2232a-7592-4e21-a432-35a2d66b6c79Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang et al.ICCV 2021 · 376 citations
Related papers
- CLIP-Count: Towards Text-Guided Zero-Shot Object CountingRuixiang Jiang, Lingbo Liu, Changwen ChenACM MM 2023 · 78 citations
- Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language ModelsJinhao Li, Haopeng Li, Sarah Monazam Erfani, Lei Feng et al.ICML 2024 · 30 citations
- VLCounter: Text-Aware Visual Representation for Zero-Shot Object CountingSeunggu Kang, WonJun Moon, Euiyeon Kim, Jae-Pil HeoAAAI 2024 · 69 citations
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based SelectionLincan Cai, Jingxuan Kang, Shuang Li, Wenxuan Ma et al.ICML 2025
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingYifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei et al.CVPR 2025
