CLIP-Count: Towards Text-Guided Zero-Shot Object Counting
Ruixiang Jiang, Lingbo Liu, Changwen Chen
摘要
Recent advances in visual-language models have shown remarkable zero-shot text-image matching ability that is transferable to downstream tasks such as object detection and segmentation. Adapting these models for object counting, however, remains a formidable challenge. In this study, we first investigate transferring vision-language models (VLMs) for class-agnostic object counting. Specifically, we propose CLIP-Count, the first end-to-end pipeline that estimates density maps for open-vocabulary objects with text guidance in a zero-shot manner. To align the text embedding with dense visual features, we introduce a patch-text contrastive loss that guides the model to learn informative patch-level visual representations for dense prediction. Moreover, we design a hierarchical patch-text interaction module to propagate semantic information across different resolution levels of visual features. Benefiting from the full exploitation of the rich image-text alignment knowledge of pretrained VLMs, our method effectively generates high-quality density maps for objects-of-interest. Extensive experiments on FSC-147, CARPK, and ShanghaiTech crowd counting datasets demonstrate state-of-the-art accuracy and generalizability of the proposed method. Code is available: https://github.com/songrise/CLIP-Count. https://github.com/songrise/CLIP-Count.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 被引用 96 次
- DAVE - A Detect-and-Verify Paradigm for Low-Shot CountingJer Pelhan, Alan Lukezic, Vitjan Zavrtanik, Matej KristanCVPR 2024 · 被引用 14 次
- OmniCount: Multi-label Object Counting with Semantic-Geometric PriorsAnindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan DuttaAAAI 2025 · 被引用 14 次
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 被引用 14 次
- A Fixed-Point Approach to Unified Prompt-Based CountingWei Lin, Antoni B. ChanAAAI 2024 · 被引用 11 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- CrowdCLIP: Unsupervised Crowd Counting via Vision-Language ModelDingkang Liang, Jiahao Xie, Zhikang Zou, Xiaoqing Ye 等CVPR 2023
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 被引用 7 次
- VLCounter: Text-Aware Visual Representation for Zero-Shot Object CountingSeunggu Kang, WonJun Moon, Euiyeon Kim, Jae-Pil HeoAAAI 2024 · 被引用 69 次
