CountGD: Multi-Modal Open-World Counting
Niki Amini-Naieni, Tengda Han, Andrew Zisserman
Abstract
The goal of this paper is to improve the generality and accuracy of open-vocabulary object counting in images. To improve the generality, we repurpose an open-vocabulary detection foundation model (GroundingDINO) for the counting task, and also extend its capabilities by introducing modules to enable specifying the target object to count by visual exemplars. In turn, these new capabilities - being able to specify the target object by multi-modalites (text and exemplars) - lead to an improvement in counting accuracy. We make three contributions: First, we introduce the first open-world counting model, CountGD, where the prompt can be specified by a text description or visual exemplars or both; Second, we show that the performance of the model significantly improves the state of the art on multiple counting benchmarks - when using text only, CountGD is comparable to or outperforms all previous text-only works, and when using both text and visual exemplars, we outperform all previous models; Third, we carry out a preliminary study into different interactions between the text and visual exemplar prompts, including the cases where they reinforce each other and where one restricts the other. The code and an app to test the model are available at https://www.robots.ox.ac.uk/ vgg/research/countgd/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40d5c3e6-d129-4ff0-b0eb-438737a38c2cCited by top-tier papers15
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 14 citations
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 7 citations
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao et al.CVPR 2026 · 6 citations
- Open-World Object Counting in VideosNiki Amini-Naieni, Andrew ZissermanAAAI 2026 · 6 citations
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object CountingAtin Pothiraj, Elias Stengel-Eskin, Jaemin Cho, Mohit BansalICCV 2025 · 4 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu et al.CVPR 2022 · 99 citations
Related papers
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 69 citations
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution ShiftsJiansheng Li, Xingxuan Zhang, Hao Zou, Yige Guo et al.CVPR 2025
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationGuanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu et al.ICCV 2025 · 4 citations
- Exploring Contextual Attribute Density in Referring Expression CountingZhicheng Wang, Zhiyu Pan, Zhan Peng, Jian Cheng et al.CVPR 2025
- OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesZehong Ma, Shiliang Zhang, Longhui Wei, Qi TianCVPR 2024 · 6 citations
