CountGD: Multi-Modal Open-World Counting
Niki Amini-Naieni, Tengda Han, Andrew Zisserman
摘要
The goal of this paper is to improve the generality and accuracy of open-vocabulary object counting in images. To improve the generality, we repurpose an open-vocabulary detection foundation model (GroundingDINO) for the counting task, and also extend its capabilities by introducing modules to enable specifying the target object to count by visual exemplars. In turn, these new capabilities - being able to specify the target object by multi-modalites (text and exemplars) - lead to an improvement in counting accuracy. We make three contributions: First, we introduce the first open-world counting model, CountGD, where the prompt can be specified by a text description or visual exemplars or both; Second, we show that the performance of the model significantly improves the state of the art on multiple counting benchmarks - when using text only, CountGD is comparable to or outperforms all previous text-only works, and when using both text and visual exemplars, we outperform all previous models; Third, we carry out a preliminary study into different interactions between the text and visual exemplar prompts, including the cases where they reinforce each other and where one restricts the other. The code and an app to test the model are available at https://www.robots.ox.ac.uk/ vgg/research/countgd/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 被引用 14 次
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 被引用 7 次
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao 等CVPR 2026 · 被引用 6 次
- Open-World Object Counting in VideosNiki Amini-Naieni, Andrew ZissermanAAAI 2026 · 被引用 6 次
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object CountingAtin Pothiraj, Elias Stengel-Eskin, Jaemin Cho, Mohit BansalICCV 2025 · 被引用 4 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu 等ICCV 2019 · 被引用 835 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu 等CVPR 2022 · 被引用 99 次
相关 Paper
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 被引用 69 次
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution ShiftsJiansheng Li, Xingxuan Zhang, Hao Zou, Yige Guo 等CVPR 2025
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationGuanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu 等ICCV 2025 · 被引用 4 次
- Exploring Contextual Attribute Density in Referring Expression CountingZhicheng Wang, Zhiyu Pan, Zhan Peng, Jian Cheng 等CVPR 2025
- OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesZehong Ma, Shiliang Zhang, Longhui Wei, Qi TianCVPR 2024 · 被引用 6 次
