TrueCount: Improving Open-World Object Counting with Visual-Language Models and Dynamic Multi-Modal Inputs
Ziqiang Shi, Rujie Liu, Jun Takahashi, Shan Jiang
摘要
Although object counting based on two-dimensional (2D) RGB images offers an effective solution in certain scenarios, it is significantly challenged by complex environments characterized by background noise, occlusion, depth variations, illumination changes, and other factors, often resulting in miscounting or missed detections. This issue is particularly pronounced in applications such as agriculture or retail, where fruits or products are frequently stacked in layers on shelves, severely compromising counting accuracy. To address these limitations, we propose TrueCount, a novel method that integrates segmented images, depth information, and other multi-modal data to enhance counting accuracy and robustness of pretrained large vision-language models (VLMs). TrueCount introduces a flexible framework capable of simultaneously processing multiple modal signals, including 2D RGB images, segmentation, and depth maps, while supporting both textual and visual prompts. During training and inference, TrueCount performs cross-attention and self-attention across all inputs and prompts. These features are then decoded to localize prompt features within the input, thereby jointly optimizing the model's counting capability. Additionally, TrueCount dynamically assesses the confidence of each modality for accurate counting in a given context, enabling effective fusion and complementary utilization of multi-modal information. Extensive experiments on multiple benchmarks including FSC-147 and CountBench demonstrate that TrueCount surpasses the previous state-of-the-art, e.g. achieving a new minimum mean average error of 4.64 on FSC-147.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 被引用 7 次
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement LearningYuqi Liu, Tianyuan Qu, Zhisheng Zhong, Bohao Peng 等ICLR 2026 · 被引用 15 次
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang 等ICCV 2025 · 被引用 1 次
- Benchmarking Dense and Indiscernible Object Counting with BlueberriesWeihao Bo, Yanpeng Sun, Jingwen Qin, Fei Shen 等ICML 2026
- UNICBench: UNIfied Counting Benchmark for MLLMChenggang Rong, Tao Han, Zhiyuan Zhao, Yaowu Fan 等CVPR 2026 · 被引用 3 次
