OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
Anindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan Dutta
Abstract
Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, these strategies come with their own set of limitations, such as the need for manual exemplar input and multiple passes for multiple categories, resulting in significant inefficiencies. This paper introduces a more practical approach enabling simultaneous counting of multiple object categories using an open-vocabulary framework. Our solution, Om-niCount, stands out by using semantic and geometric insights (priors) from pre-trained models to count multiple categories of objects as specified by users, all without additional training. OmniCount distinguishes itself by generating precise object masks and leveraging varied interactive prompts via the Segment Anything Model for efficient counting. To evaluate OmniCount, we created the OmniCount-191 benchmark, a first-of-its-kind dataset with multi-label object counts, including points, bounding boxes, and VQA annotations. Our comprehensive evaluation in OmniCount-191, alongside other leading benchmarks, demonstrates Om-niCount's exceptional performance, significantly outpacing existing solutions. The project webpage is available at https://mondalanindya.github.io/OmniCount .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 14 citations
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual ReasoningZejun Li, Yingxiu Zhao, Jiwen Zhang, Siyuan Wang et al.ICLR 2026 · 7 citations
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao et al.CVPR 2026 · 6 citations
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
Related papers
- Open-World Object Counting in VideosNiki Amini-Naieni, Andrew ZissermanAAAI 2026 · 6 citations
- OmniLabel: A Challenging Benchmark for Language-Based Object DetectionSamuel Schulter, Vijay Kumar B. G, Yumin Suh, Konstantinos M. Dafnis et al.ICCV 2023 · 18 citations
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 96 citations
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu et al.ICCV 2023 · 56 citations
- Rethinking Object Detection in Retail StoresYuanqiang Cai, Longyin Wen, Libo Zhang, Dawei Du et al.AAAI 2021 · 32 citations
