T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting
Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound
摘要
Zero-shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T2ICount, a diffusion-based framework that leverages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step denoising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines textimage feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the denoising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/cha15yq/ T2ICount .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao 等CVPR 2026 · 被引用 6 次
- Bootstrapping MLLM for Weakly‑Supervised Class‑Agnostic Object CountingXiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao 等ICLR 2026 · 被引用 3 次
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
- Ψ-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score ModelsTaehoon Yoon, Yunhong Min, Kyeongmin Yeo, Minhyuk SungNeurIPS 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu 等CVPR 2022 · 被引用 99 次
相关 Paper
- CLIP-Count: Towards Text-Guided Zero-Shot Object CountingRuixiang Jiang, Lingbo Liu, Changwen ChenACM MM 2023 · 被引用 78 次
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 被引用 7 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- VLCounter: Text-Aware Visual Representation for Zero-Shot Object CountingSeunggu Kang, WonJun Moon, Euiyeon Kim, Jae-Pil HeoAAAI 2024 · 被引用 69 次
- R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image GenerationJiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang 等ICLR 2024 · 被引用 31 次
