Embodied Crowd Counting
Runling Long, Yunlong Wang, Jia Wan, Xiang Deng, Xinting Zhu, Weili Guan, Antoni B. Chan, Liqiang Nie
Abstract
Occlusion is one of the fundamental challenges in crowd counting. In the community, various data-driven approaches have been developed to address this issue, yet their effectiveness is limited. This is mainly because most existing crowd counting datasets on which the methods are trained are based on passive cameras, restricting their ability to fully sense the environment. Recently, embodied navigation methods have shown significant potential in precise object detection in interactive scenes. These methods incorporate active camera settings, holding promise in addressing the fundamental issues in crowd counting. However, most existing methods are designed for indoor navigation, showing unknown performance in analyzing complex object distribution in large-scale scenes, such as crowds. Besides, most existing embodied navigation datasets are indoor scenes with limited scale and object quantity, preventing them from being introduced into dense crowd analysis. Based on this, a novel task, Embodied Crowd Counting (ECC), is proposed to count the number of persons in a large-scale scene actively. We then build up an interactive simulator, the Embodied Crowd Counting Dataset (ECCD), which enables large-scale scenes and large object quantities. A prior probability distribution approximating a realistic crowd distribution is introduced to generate crowds. Then, a zero-shot navigation method (ZECC) is proposed as a baseline. This method contains an MLLM-driven coarse-to-fine navigation mechanism, enabling active Z-axis exploration, and a normal-line-based crowd distribution analysis method for fine counting. Experimental results show that the proposed method achieves the best trade-off between counting accuracy and navigation cost. Code can be found at https://github.com/longrunling/ECC?.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1308619e-5925-413f-a3d5-2929d1c3238fBuilds on17
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman et al.NeurIPS 2022 · 344 citations
- RIO: 3D Object Instance Re-Localization in Changing Indoor EnvironmentsJohanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari et al.ICCV 2019 · 233 citations
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang et al.ICCV 2023 · 132 citations
- Point-Query Quadtree for Crowd Counting, Localization, and MoreChengxin Liu, Hao Lu, Zhiguo Cao, Tongliang LiuICCV 2023 · 89 citations
Related papers
- Explore and Tell: Embodied Visual Captioning in 3D EnvironmentsAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinICCV 2023 · 4 citations
- World2Minecraft: Occupancy-Driven Simulated Scenes ConstructionLechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang et al.ICLR 2026 · 1 citation
- RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured EnvironmentsHaisheng Su, Feixiang Song, Cong Ma, Wei Wu et al.CVPR 2025
- Embodied Amodal Recognition: Learning to Move to Perceive ObjectsJianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen et al.ICCV 2019 · 70 citations
- Learning from Human Gaze: Human-like Robot Social Navigation in Dense CrowdsZhecheng Yu, Yan Lyu, Chen Yang, Tao Chen et al.AAAI 2026
