Rethinking Object Detection in Retail Stores
Yuanqiang Cai, Longyin Wen, Libo Zhang, Dawei Du, Weiqiang Wang
Abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50, 394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-toend manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9258a94-00e7-44fd-b0f5-2c606875e2a2Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Kaputt: A Large-Scale Dataset for Visual Defect DetectionSebastian Höfer, Dorian Fritz Henning, Artemij Amiranashvili, Douglas Morrison et al.ICCV 2025 · 2 citations
- OmniCount: Multi-label Object Counting with Semantic-Geometric PriorsAnindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan DuttaAAAI 2025 · 14 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Learning To Count EverythingViresh Ranjan, Udbhav Sharma, Thu Nguyen, Minh HoaiCVPR 2021
- Referring Expression CountingSiyang Dai, Jun Liu, Ngai-Man CheungCVPR 2024
