Benchmarking Dense and Indiscernible Object Counting with Blueberries
Weihao Bo, Yanpeng Sun, Jingwen Qin, Fei Shen, Xiaofan Li, Zechao Li
摘要
Real-world agricultural counting often operates in the extreme regime of Dense and Indiscernible Object Counting (DIOC), where targets are tiny, clustered, and highly camouflaged. To facilitate research in this domain, we introduce DIOCblueberry, a large-scale benchmark that pushes the boundaries of visual perception. Unlike general datasets with salient objects, DIOCblueberry features extreme occlusion and camouflage. Compared to the popular FSC147 benchmark, it contains 1.9× more instances per image (avg. 108) with an average box pixel ratio that is 7.9× smaller, serving as a rigorous testbed for model robustness. Standard counting methods struggle in these scenarios due to severe visual ambiguity and scale mismatch. To address this, we propose MaskCount, a coarse-to-fine framework that incorporates semantic guidance. MaskCount leverages Vision-Language Models (CLIP) to generate pseudo segmentation masks for background suppression and employs a contrastive loss to maximize feature discriminability between fruits and foliage. Additionally, we design an edge-aware cropping mechanism to resolve boundary truncation in dense clusters. Extensive experiments demonstrate that MaskCount achieves a new state-of-the-art, reducing MAE and RMSE by 49.16% and 70.50% respectively on DIOCblueberry, with strong generalization to other agricultural scenes. Our DIOCblueberry benchmark is publicly available at https://huggingface.co/datasets/ weihao-bo/DIOCblueberry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang 等ICCV 2021 · 被引用 376 次
- Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuningYanpeng Sun, Qiang Chen, Xiangyu He, Jian Wang 等NeurIPS 2022 · 被引用 97 次
- A Low-Shot Object Counting Network With Iterative Prototype AdaptationNikola Ðukic, Alan Lukezic, Vitjan Zavrtanik, Matej KristanICCV 2023 · 被引用 91 次
相关 Paper
- CrowdCLIP: Unsupervised Crowd Counting via Vision-Language ModelDingkang Liang, Jiahao Xie, Zhikang Zou, Xiaoqing Ye 等CVPR 2023
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- TrueCount: Improving Open-World Object Counting with Visual-Language Models and Dynamic Multi-Modal InputsZiqiang Shi, Rujie Liu, Jun Takahashi, Shan JiangACM MM 2025 · 被引用 1 次
- Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsHeeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho 等NeurIPS 2024 · 被引用 32 次
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingYifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei 等CVPR 2025
