Benchmarking Dense and Indiscernible Object Counting with Blueberries
Weihao Bo, Yanpeng Sun, Jingwen Qin, Fei Shen, Xiaofan Li, Zechao Li
Abstract
Real-world agricultural counting often operates in the extreme regime of Dense and Indiscernible Object Counting (DIOC), where targets are tiny, clustered, and highly camouflaged. To facilitate research in this domain, we introduce DIOCblueberry, a large-scale benchmark that pushes the boundaries of visual perception. Unlike general datasets with salient objects, DIOCblueberry features extreme occlusion and camouflage. Compared to the popular FSC147 benchmark, it contains 1.9× more instances per image (avg. 108) with an average box pixel ratio that is 7.9× smaller, serving as a rigorous testbed for model robustness. Standard counting methods struggle in these scenarios due to severe visual ambiguity and scale mismatch. To address this, we propose MaskCount, a coarse-to-fine framework that incorporates semantic guidance. MaskCount leverages Vision-Language Models (CLIP) to generate pseudo segmentation masks for background suppression and employs a contrastive loss to maximize feature discriminability between fruits and foliage. Additionally, we design an edge-aware cropping mechanism to resolve boundary truncation in dense clusters. Extensive experiments demonstrate that MaskCount achieves a new state-of-the-art, reducing MAE and RMSE by 49.16% and 70.50% respectively on DIOCblueberry, with strong generalization to other agricultural scenes. Our DIOCblueberry benchmark is publicly available at https://huggingface.co/datasets/ weihao-bo/DIOCblueberry.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afa6a3db-80dc-4a6b-8d9f-24f1cb0a202cBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang et al.ICCV 2021 · 376 citations
- Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuningYanpeng Sun, Qiang Chen, Xiangyu He, Jian Wang et al.NeurIPS 2022 · 97 citations
- A Low-Shot Object Counting Network With Iterative Prototype AdaptationNikola Ðukic, Alan Lukezic, Vitjan Zavrtanik, Matej KristanICCV 2023 · 91 citations
Related papers
- CrowdCLIP: Unsupervised Crowd Counting via Vision-Language ModelDingkang Liang, Jiahao Xie, Zhikang Zou, Xiaoqing Ye et al.CVPR 2023
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- TrueCount: Improving Open-World Object Counting with Visual-Language Models and Dynamic Multi-Modal InputsZiqiang Shi, Rujie Liu, Jun Takahashi, Shan JiangACM MM 2025 · 1 citation
- Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsHeeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho et al.NeurIPS 2024 · 32 citations
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingYifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei et al.CVPR 2025
