Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
Da Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao
Abstract
Zero-shot object counting (ZSOC) aims to enumerate objects of arbitrary categories specified by text descriptions without requiring visual exemplars. However, existing methods often treat counting as a coarse retrieval task, suffering from a lack of fine-grained quantity awareness. Furthermore, they frequently exhibit spatial insensitivity and degraded generalization due to feature space distortion during model adaptation.To address these challenges, we present QICA, a novel framework that synergizes quantity perception with robust spatial cast aggregation. Specifically, we introduce a Synergistic Prompting Strategy (SPS) that adapts vision and language encoders through numerically conditioned prompts, bridging the gap between semantic recognition and quantitative reasoning. To mitigate feature distortion, we propose a Cost Aggregation Decoder (CAD) that operates directly on vision-text similarity maps. By refining these maps through spatial aggregation, CAD prevents overfitting while preserving zero-shot transferability. Additionally, a multi-level quantity alignment loss () is employed to enforce numerical consistency across the entire pipeline. Extensive experiments on FSC-147 demonstrate competitive performance, while zero-shot evaluation on CARPK and ShanghaiTech-A validates superior generalization to unseen domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67cf5249-5b45-4fb0-a6b4-080b56d7e8c7Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- Rethinking Spatial Invariance of Convolutional Networks for Object CountingZhi-Qi Cheng, Qi Dai, Hong Li, Jingkuan Song et al.CVPR 2022 · 119 citations
Related papers
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 7 citations
- A Low-Shot Object Counting Network With Iterative Prototype AdaptationNikola Ðukic, Alan Lukezic, Vitjan Zavrtanik, Matej KristanICCV 2023 · 91 citations
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingYifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei et al.CVPR 2025
- CountSE: Soft Exemplar Open-Set Object CountingShuai Liu, Peng Zhang, Shiwei Zhang, Wei KeICCV 2025 · 1 citation
- VLCounter: Text-Aware Visual Representation for Zero-Shot Object CountingSeunggu Kang, WonJun Moon, Euiyeon Kim, Jae-Pil HeoAAAI 2024 · 69 citations
